CloudLinux

CloudLinux

CloudLinux is a leading provider of the CloudLinux OS, a platform for Linux web hosting that offers next-level performance and security. With a focus on optimizing web hosting environments, CloudLinux helps service providers improve density, stability,...

IT Services
51-250
Founded 2009

Description

  • Define SLIs, SLOs, error budgets, tiers, and ownership for approximately 70 components.
  • Develop service, fleet, control-efficacy, delivery, and pipeline measurement standards.
  • Design and build a privacy-conscious telemetry collection pipeline for customer fleets and cloud services.
  • Extend agent and service instrumentation in Python, Go, and Rust.
  • Consolidate dashboards, queries, and reporting into maintainable observability tools.
  • Implement symptom-based, SLO-driven alerting with burn-rate semantics and page/ticket/dashboard tiers.
  • Create alert ownership, runbooks, failure-mode documentation, and ongoing alert-hygiene practices.
  • Build machine-readable ownership maps, severity policies, acknowledgement SLAs, and follow-the-sun escalation routing.
  • Establish incident command practices, blameless postmortems, and action-item follow-through.
  • Coach engineering squads to own their pagers and operational practices while operating the observability platform.

Requirements

  • Substantial production engineering or SRE experience, including defining an SLO framework from scratch.
  • Strong Python skills and ability to read and modify Go or Rust.
  • Practical experience with Prometheus/OpenMetrics, Grafana, Alertmanager-class routing, and ClickHouse or an equivalent columnar store.
  • Experience debugging distributed systems on bare metal and long-lived hosts outside Kubernetes.
  • Production-scale configuration management and CI experience with Ansible, GitLab CI, Jenkins, or equivalents.
  • Experience designing push telemetry, sampling, clock-skew handling, partial reporting, and privacy controls for machines not directly owned or scrapeable.
  • Strong asynchronous written communication and cross-team facilitation skills.
  • Security product experience with WAF, EDR, antivirus, or vulnerability management (preferred).
  • Experience with audit-oriented monitoring, such as SOC 2, ISO 27001, or NIST continuous monitoring (preferred).
  • Experience with OpenTelemetry, eBPF, Sentry, cost- and cardinality-aware telemetry, agentic development tooling, or Kubernetes (preferred).

Benefits

  • Fully remote work with flexible hours from anywhere worldwide.
  • 24 paid vacation days, 10 national holidays, and unlimited sick leave.
  • Private medical insurance compensation.
  • Co-working and gym or sports reimbursement.
  • Professional development through challenging projects, mentoring, and knowledge exchange.
  • Opportunity to receive a reward for an innovative patentable idea.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Site Reliability Engineer II

GoGuardian 251-1K Internet Software & Services

GoGuardian is seeking a Site Reliability Engineer II to build, maintain, and scale the cloud infrastructure, data services, and developer tooling that support reliable and secure K-12 learning products.

AWS CI/CD GitHub Actions Go JavaScript Jenkins Kubernetes Linux MongoDB OpenSearch Python Serverless Terraform TypeScript
1 day, 6 hours ago

Senior Site Reliability Engineer

Megaport 251-1K Diversified Telecommunication Services

Megaport is seeking a Senior Platform Engineer to strengthen DevOps and SRE practices, improve the reliability and security of production systems, and support customer and company outcomes across a globally distributed team.

AWS Bash Cassandra CI/CD ClickHouse Git GitHub Go Kubernetes Linux PostgreSQL Python Terraform
1 day, 7 hours ago

Site Reliability Engineer, Tech Lead

Loadsmart 251-1K Air Freight & Logistics

Loadsmart is hiring a remote SRE Tech Lead in Brazil to build and operate its internal engineering platform, improve reliability, and enable safe, dependable applications across engineering teams.

Ansible AWS Bash Chef CI/CD Docker Kubernetes PostgreSQL Python Terraform
2 days, 7 hours ago

Senior Site Reliability Engineer (Performance and Scalability)

Digitalzone 251-1K Media

DigitalZone is hiring an SRE/platform engineer to build the scalability, observability, and resilience foundations that help engineering teams handle large campaign traffic spikes reliably.

AWS Go Laravel PHP PostgreSQL TypeScript
3 days, 7 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers