Honeycomb.io

Honeycomb.io

Honeycomb.io provides a comprehensive observability platform designed for engineers to effectively debug and monitor distributed services, including microservices and serverless applications, facilitating collaborative problem-solving and enhancing ove...

Internet Software & Services
51-250
Founded 2016
$149M raised

Description

  • Define architecture and operational standards for Refinery as a Service and Honeycomb Private Cloud across multiple AWS accounts and regions.
  • Architect Terraform modules, Helm charts, and deployment automation for the FRE team.
  • Own capacity planning, scaling strategy, upgrade sequencing, and cost optimization for managed infrastructure.
  • Serve as the final technical escalation point for novel, high-stakes customer issues.
  • Resolve deep infrastructure and observability problems across distributed systems, Kubernetes, AWS networking, and service meshes.
  • Partner with customer SRE, platform, and engineering leaders on escalations and architecture redesigns.
  • Provide senior incident command for managed services and build playbooks and diagnostic tooling for other engineers.
  • Shape Honeycomb’s open source strategy in the OpenTelemetry ecosystem and contribute to Refinery and the Honeycomb Collector Distro.
  • Lead architecture reviews, SLO workshops, instrumentation deep-dives, and high-stakes POCs and pilots.
  • Build internal tools and UIs, mentor engineers, and drive cross-functional alignment across field, product, support, and engineering.

Requirements

  • 9+ years of experience in engineering, SRE, infrastructure, DevOps, or equivalent with staff-level scope and impact.
  • Deep hands-on experience with Kubernetes, with EKS strongly preferred.
  • Strong AWS expertise across EC2, EKS, ECS, ALB/NLB, VPC, PrivateLink, IAM, S3, and Route53.
  • Experience with multi-account architecture design, service quotas, and cost optimization.
  • A track record of senior incident command and ownership of incident response and postmortem improvements.
  • Infrastructure as Code mastery with Terraform, Helm, Chef, or Ansible.
  • Deep observability expertise including logging, tracing, metrics, SLOs/SLIs, and instrumentation lifecycle standards.
  • Strong command of OpenTelemetry or equivalent, including community contribution and leadership.
  • Proficiency in at least two of Go, Python, Java, TypeScript/Node.js, or .NET.
  • Excellent executive communication skills and the ability to lead in ambiguous, high-pressure situations.
  • Background in customer-facing engineering functions such as solutions architecture, field engineering, or technical consulting (nice to have).
  • Public recognition or leadership in the CNCF/OpenTelemetry ecosystem, such as maintainer status, SIG leadership, or speaking (nice to have).
  • Experience operating telemetry pipelines, managed SaaS deployments, private cloud offerings, or multi-tenant infrastructure (nice to have).
  • Prior experience at an observability, monitoring, or developer tools vendor (nice to have).

Benefits

  • On-target earnings of $200,000-$240,000 USD based on level of experience (base + commission).
  • Generous equity with an employee-friendly stock program.
  • Transparent pay based on levels relative to experience.
  • Unlimited PTO.
  • Distributed-first, remote-friendly work culture.
  • Home office, co-working, and internet stipend.
  • Full benefits coverage for employees, with additional coverage available for dependents.
  • Up to 16 weeks of paid parental leave.
  • Annual development allowance.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Vice President, Global Production Operations & Reliability

Everbridge 1K-5K Internet Software & Services

Everbridge is hiring a Vice President, Global Production Operations & Reliability to lead the company’s global production operations for its cloud-native SaaS platform and drive reliability, scalability, security, and operational excellence.

AWS CI/CD Kubernetes
20 hours, 12 minutes ago

DevOps Engineer - SRE Observability

Lingaro 5K-10K IT Services

An infrastructure-focused role at Lingaro responsible for monitoring, automating, and designing cloud systems within an Azure-based environment.

Azure Azure Pipelines CI/CD Docker GitHub GitHub Actions Grafana Kubernetes MySQL PostgreSQL Prometheus SQL Terraform
20 hours, 12 minutes ago

Site Reliability Engineer

Yuno 51-200 Payment Processing Software

Yuno is seeking a Staff Site Reliability Engineer to define and lead reliability for its AWS-based platform that provisions and manages AI agents powering global payments at scale.

Apache Airflow AWS Databricks Datadog Docker EC2 Fly.io GCP Go Kafka Kubernetes MLflow MLOps MongoDB NATS OpsGenie PagerDuty PostgreSQL Prefect Pulumi Python RabbitMQ Railway Redis Snowflake SQL Terraform
20 hours, 57 minutes ago

Associate Technical Account Manager

ActivTrak 51-250 Professional Services

ActivTrak is hiring a remote Associate Technical Account Manager to support strategic customers by managing technical needs, resolving account issues, and helping customers maximize value from the platform.

GCP macOS Power BI Python SQL Tableau
20 hours, 57 minutes ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers