Senior Site Reliability Engineer

2 months, 1 week ago
Full-time
Senior
DevOps and Infrastructure
Honeycomb.io

Honeycomb.io

Honeycomb.io provides a comprehensive observability platform designed for engineers to effectively debug and monitor distributed services, including microservices and serverless applications, facilitating collaborative problem-solving and enhancing ove...

Internet Software & Services
51-250
Founded 2016
$149M raised

Description

  • Help scale backend systems to support Honeycomb’s highest-volume customers.
  • Work with backend teams to analyze and optimize infrastructure and the broader stack.
  • Build organizational trust through transparent communication and direct, kind feedback.
  • Train as an Incident Commander and help train others in the role.
  • Support and help develop a healthy cross-Atlantic engineering culture.
  • Participate in the EU side of the team’s follow-the-sun on-call rotation.
  • Help the organization balance reliability with other business goals and priorities.
  • Optionally represent Honeycomb externally through blog posts, conference talks, and presentations with DevRel support.

Requirements

  • Strong experience in AWS and Kubernetes.
  • Experience performing cost analysis and cost reduction.
  • Solid experience with Helm, Terraform, and CI/CD.
  • Project management skills.
  • Software engineering experience; Golang is a plus.
  • Performance engineering experience is a plus.
  • Experience with Kafka or another high-volume distributed system.
  • Excellent written and spoken communication skills, including tailoring communication to the audience and giving direct feedback.
  • Familiarity with observability concepts such as SLOs and instrumentation, plus data-driven decision making.
  • Comfort operating in ambiguity with a bias for action and experimentation.
  • Interest in both the technical and human sides of reliability engineering.
  • Experience working in geographically distributed teams.
  • Please note that Honeycomb cannot currently sponsor or support visa transfers.
  • All hires must verify identity and eligibility to work.

Benefits

  • Base salary of €140,590 to €165,400 EUR depending on experience.
  • Generous equity with an employee-friendly stock program.
  • Transparent pay levels based on experience.
  • Unlimited PTO.
  • Home office, co-working, and internet stipend.
  • Full benefits coverage for employees, with additional coverage available for dependents.
  • Up to 16 weeks of paid parental leave, regardless of path to parenthood.
  • Annual development allowance.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Vice President, Global Production Operations & Reliability

Everbridge 1K-5K Internet Software & Services

Everbridge is hiring a Vice President, Global Production Operations & Reliability to lead the company’s global production operations for its cloud-native SaaS platform and drive reliability, scalability, security, and operational excellence.

AWS CI/CD Kubernetes
20 hours, 25 minutes ago

DevOps Engineer - SRE Observability

Lingaro 5K-10K IT Services

An infrastructure-focused role at Lingaro responsible for monitoring, automating, and designing cloud systems within an Azure-based environment.

Azure Azure Pipelines CI/CD Docker GitHub GitHub Actions Grafana Kubernetes MySQL PostgreSQL Prometheus SQL Terraform
20 hours, 25 minutes ago

Site Reliability Engineer

Yuno 51-200 Payment Processing Software

Yuno is seeking a Staff Site Reliability Engineer to define and lead reliability for its AWS-based platform that provisions and manages AI agents powering global payments at scale.

Apache Airflow AWS Databricks Datadog Docker EC2 Fly.io GCP Go Kafka Kubernetes MLflow MLOps MongoDB NATS OpsGenie PagerDuty PostgreSQL Prefect Pulumi Python RabbitMQ Railway Redis Snowflake SQL Terraform
21 hours, 10 minutes ago

Senior Site Reliability Engineer

PandaDoc 251-1K Internet Software & Services

PandaDoc is hiring a Site Reliability Engineer to help keep its document workflow platform reliable, resilient, and available for customers.

Agile AWS Django Grafana Java Kafka Kubernetes NATS PostgreSQL Python RabbitMQ Spring Boot
1 day, 20 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers