CloudLinux

CloudLinux

CloudLinux is a leading provider of the CloudLinux OS, a platform for Linux web hosting that offers next-level performance and security. With a focus on optimizing web hosting environments, CloudLinux helps service providers improve density, stability,...

IT Services
51-250
Founded 2009

Description

  • Operate the observability platform, including team onboarding, alerting, cost, and capacity management.
  • Manage GitLab and the CI runner fleet, including upgrades, access, capacity, backups, and restore drills.
  • Maintain platform services with appropriate monitoring, documentation, and runbooks.
  • Research, design, and deploy new services as code with monitoring, backups, and documentation.
  • Support developer requests involving access, onboarding, pipelines, exporters, and dashboards, converting recurring needs into self-service.
  • Lead incident response, service restoration, root-cause analysis, post-mortems, and prevention improvements.
  • Deliver infrastructure changes through reviewed merge requests with appropriate planning and validation.
  • Write runbooks, onboarding guides, maintenance notices, and status updates for engineers.
  • Use AI engineering assistants to delegate scoped work, review outputs, and document lessons learned.

Requirements

  • Senior-level infrastructure, platform, or site reliability engineering experience, including ownership of at least one production service.
  • Linux administration and debugging on bare metal and virtual machines.
  • Production Kubernetes experience delivered through GitOps, including performing cluster upgrades.
  • Experience with infrastructure as code using Ansible and Terraform or OpenTofu, with merge-request reviews.
  • Production GitLab administration and GitLab CI experience, or equivalent depth with another CI system.
  • Working knowledge of Prometheus and Grafana, including alert rules, dashboards, and PromQL.
  • Ability to write clear technical documentation for engineers outside the team.
  • Strong communication skills and the ability to manage scope, priorities, expectations, and stakeholder pushback.
  • Advanced experience with AI engineering assistants such as Claude or Codex, including agent loops, validation, testing, and production safeguards.
  • Upper-intermediate or higher English proficiency.
  • Experience with SLOs and burn-rate alerting, CI microVM isolation, S3-compatible storage, AWS cost management, or Kafka/ClickHouse/Redis-backed systems is preferred.
  • Python or Go experience for exporters and internal services is preferred.

Benefits

  • Fully remote work with flexible hours worldwide.
  • 24 paid vacation days, 10 national holidays, and unlimited sick leave.
  • Compensation for private medical insurance.
  • Co-working and gym or sports reimbursement.
  • Professional development and education budget.
  • Interesting and challenging projects.
  • Opportunity to receive a reward for an innovative patentable idea.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Principal Software Engineer - Builder Experience - Platform Engineering Productivity

Elastic 1K-5K Internet Software & Services

Elastic is hiring a Principal Software Engineer for its Platform Engineering Productivity team to improve the systems, tooling, and practices used to build, test, and release products at scale.

Bash CI/CD Docker Git Go Java Kubernetes Python Scala Terraform
1 day, 20 hours ago

CX Tooling Engineer

Spring Health 1K-5K Health Care Providers & Services

Spring Health is hiring a remote CX Tooling Engineer to design, operate, and improve the customer experience and support systems that enable reliable, scalable care operations in a regulated environment.

JSON LLM REST API
1 day, 20 hours ago

Principal Software Engineer - Builder Experience - Platform Engineering Productivity

Elastic 1K-5K Internet Software & Services

Elastic is hiring a Principal Software Engineer for its Platform Engineering Productivity team to improve the systems, tooling, and practices that enable reliable, efficient software delivery at scale.

Bash CI/CD Docker Git Go Java Kubernetes Python Scala Terraform
1 day, 21 hours ago

Principal Software Engineer - Builder Experience - Platform Engineering Productivity

Elastic 1K-5K Internet Software & Services

Elastic is hiring a Principal Software Engineer for its Platform Engineering Productivity team to improve the systems, tooling, and practices that help distributed engineering teams build, test, and release products efficiently and reliably.

Bash CI/CD Docker Git Go Java Kubernetes Python Scala System Design Terraform
1 day, 21 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers