Senior Site Reliability Engineer

1 month, 2 weeks ago
Full-time
Senior
DevOps and Infrastructure
Honeycomb.io

Honeycomb.io

Honeycomb.io provides a comprehensive observability platform designed for engineers to effectively debug and monitor distributed services, including microservices and serverless applications, facilitating collaborative problem-solving and enhancing ove...

Internet Software & Services
51-250
Founded 2016
$149M raised

Description

  • Help scale backend systems to support Honeycomb’s highest-volume customers.
  • Work with backend teams to analyze and optimize infrastructure and the broader stack.
  • Build organizational trust through transparent communication and direct, kind feedback.
  • Train as an Incident Commander and help train others in the role.
  • Support and help develop a healthy cross-Atlantic engineering culture.
  • Participate in the EU side of the team’s follow-the-sun on-call rotation.
  • Help the organization balance reliability with other business goals and priorities.
  • Optionally represent Honeycomb externally through blog posts, conference talks, and presentations with DevRel support.

Requirements

  • Strong experience in AWS and Kubernetes.
  • Experience performing cost analysis and cost reduction.
  • Solid experience with Helm, Terraform, and CI/CD.
  • Project management skills.
  • Software engineering experience; Golang is a plus.
  • Performance engineering experience is a plus.
  • Experience with Kafka or another high-volume distributed system.
  • Excellent written and spoken communication skills, including tailoring communication to the audience and giving direct feedback.
  • Familiarity with observability concepts such as SLOs and instrumentation, plus data-driven decision making.
  • Comfort operating in ambiguity with a bias for action and experimentation.
  • Interest in both the technical and human sides of reliability engineering.
  • Experience working in geographically distributed teams.
  • Please note that Honeycomb cannot currently sponsor or support visa transfers.
  • All hires must verify identity and eligibility to work.

Benefits

  • Base salary of €140,590 to €165,400 EUR depending on experience.
  • Generous equity with an employee-friendly stock program.
  • Transparent pay levels based on experience.
  • Unlimited PTO.
  • Home office, co-working, and internet stipend.
  • Full benefits coverage for employees, with additional coverage available for dependents.
  • Up to 16 weeks of paid parental leave, regardless of path to parenthood.
  • Annual development allowance.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

Counterpart Health 51-200 hospital & health care

Counterpart Health is hiring a Senior Site Reliability and Infrastructure Engineer to support and evolve the technology platform behind its primary care tool and maintain reliable infrastructure for domestic and international workloads.

AWS Azure CI/CD Containerd DNS Docker GCP Go gRPC Helm Kubernetes Linux Load Balancing Prometheus Python Shell Scripting TCP/IP
15 hours, 34 minutes ago

Senior Test Platform & Reliability Engineer - Star Trek Fleet Command

Scopely 1K-5K Internet Software & Services

Scopely is hiring a Senior Test Platform & Reliability Engineer in Ireland to build validation, reliability, and developer enablement platforms for Star Trek Fleet Command’s large-scale live-service backend systems.

AWS Bash CI/CD Docker GitLab Go Python Terraform
15 hours, 49 minutes ago

Senior Software Engineer - Databases, SRE | Canada | Remote

Grafana 1K-5K IT Services

Grafana Labs is hiring a Senior Software Engineer for its remote SRE team to improve reliability and operability of Grafana Cloud database services for high-SLA customers across AWS, GCP, and Azure.

AWS Azure GCP Go Helm Java Kubernetes Linux Microservices Python Terraform
1 day, 14 hours ago

Senior Site Reliability Engineer

Semios 51-250 Food Products

Semios Group is hiring a Senior Site Reliability Engineer to help scale, secure, and improve the reliability of its global agricultural technology platform.

AWS Azure Bash Buildkite CI/CD Datadog Docker Envoy GCP Git GitHub GitHub Actions GitLab Go Jenkins Kubernetes Linux NATS New Relic Prometheus Python Ruby Splunk Terraform
1 day, 16 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers