Honeycomb.io

Honeycomb.io

Honeycomb.io provides a comprehensive observability platform designed for engineers to effectively debug and monitor distributed services, including microservices and serverless applications, facilitating collaborative problem-solving and enhancing ove...

Internet Software & Services
51-250
Founded 2016
$149M raised

Description

  • Define architecture and operational standards for Refinery as a Service and Honeycomb Private Cloud across multiple AWS accounts and regions.
  • Architect Terraform modules, Helm charts, and deployment automation for the FRE team.
  • Own capacity planning, scaling strategy, upgrade sequencing, and cost optimization for managed infrastructure.
  • Serve as the final technical escalation point for novel, high-stakes customer issues.
  • Resolve deep infrastructure and observability problems across distributed systems, Kubernetes, AWS networking, and service meshes.
  • Partner with customer SRE, platform, and engineering leaders on escalations and architecture redesigns.
  • Provide senior incident command for managed services and build playbooks and diagnostic tooling for other engineers.
  • Shape Honeycomb’s open source strategy in the OpenTelemetry ecosystem and contribute to Refinery and the Honeycomb Collector Distro.
  • Lead architecture reviews, SLO workshops, instrumentation deep-dives, and high-stakes POCs and pilots.
  • Build internal tools and UIs, mentor engineers, and drive cross-functional alignment across field, product, support, and engineering.

Requirements

  • 9+ years of experience in engineering, SRE, infrastructure, DevOps, or equivalent with staff-level scope and impact.
  • Deep hands-on experience with Kubernetes, with EKS strongly preferred.
  • Strong AWS expertise across EC2, EKS, ECS, ALB/NLB, VPC, PrivateLink, IAM, S3, and Route53.
  • Experience with multi-account architecture design, service quotas, and cost optimization.
  • A track record of senior incident command and ownership of incident response and postmortem improvements.
  • Infrastructure as Code mastery with Terraform, Helm, Chef, or Ansible.
  • Deep observability expertise including logging, tracing, metrics, SLOs/SLIs, and instrumentation lifecycle standards.
  • Strong command of OpenTelemetry or equivalent, including community contribution and leadership.
  • Proficiency in at least two of Go, Python, Java, TypeScript/Node.js, or .NET.
  • Excellent executive communication skills and the ability to lead in ambiguous, high-pressure situations.
  • Background in customer-facing engineering functions such as solutions architecture, field engineering, or technical consulting (nice to have).
  • Public recognition or leadership in the CNCF/OpenTelemetry ecosystem, such as maintainer status, SIG leadership, or speaking (nice to have).
  • Experience operating telemetry pipelines, managed SaaS deployments, private cloud offerings, or multi-tenant infrastructure (nice to have).
  • Prior experience at an observability, monitoring, or developer tools vendor (nice to have).

Benefits

  • On-target earnings of $200,000-$240,000 USD based on level of experience (base + commission).
  • Generous equity with an employee-friendly stock program.
  • Transparent pay based on levels relative to experience.
  • Unlimited PTO.
  • Distributed-first, remote-friendly work culture.
  • Home office, co-working, and internet stipend.
  • Full benefits coverage for employees, with additional coverage available for dependents.
  • Up to 16 weeks of paid parental leave.
  • Annual development allowance.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Technical Account Manager

teamified.com Hotels, Restaurants & Leisure

Technical Account Manager supporting a global SaaS volunteer-engagement platform by advising strategic customers, solving technical challenges, and driving adoption, retention, expansion, and long-term success.

10 hours, 17 minutes ago

Senior Site Reliability Specialist II

Everbridge 1K-5K Internet Software & Services

Everbridge is hiring a Senior Site Reliability Engineer II to build resilient platforms and improve the reliability, scalability, and operational readiness of systems supporting critical-event communications.

CI/CD Kubernetes Linux
1 day, 10 hours ago

Site Reliability Engineer (SRE)

Rocket.net 11-50 IT Services

Rocket.net is seeking a Site Reliability Engineer to maintain the reliability and performance of its hosting platform while resolving complex infrastructure issues and providing advanced support to customers.

Apache Bash CDN Cloudflare Datadog DNS Linux MariaDB MySQL Nginx Redis SSH WAF WordPress
1 day, 10 hours ago

Senior Technical Account Manager

Airship 251-1K Internet Software & Services

Airship is hiring a senior Technical Account Manager to advise strategic enterprise clients, lead complex integrations and escalations, and maximize adoption of its cross-channel customer experience platform.

Android Digital Marketing Email Marketing iOS Java Linux macOS Python Swift
2 days, 9 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers