CloudLinux

CloudLinux

CloudLinux is a leading provider of the CloudLinux OS, a platform for Linux web hosting that offers next-level performance and security. With a focus on optimizing web hosting environments, CloudLinux helps service providers improve density, stability,...

IT Services
51-250
Founded 2009

Description

  • Design, build, and operate automated protection pipelines from threat intelligence through validated rule deployment.
  • Develop progressive release automation with staged rollouts, automated quality gates, holds, and rollbacks.
  • Transform multi-stage jobs into resumable, idempotent, observable systems with explicit state and recovery paths.
  • Define and enforce latency budgets and SLOs for each pipeline stage.
  • Build metrics, dashboards, alerting, and health gates that identify pipeline issues automatically.
  • Implement blast-radius controls, kill switches, safe defaults, and graceful handling of dependency failures.
  • Reduce manual operations and improve system maintainability and reliability.
  • Write unit and integration tests covering concurrency, partial failures, external API issues, and state management.
  • Investigate issues across ClickHouse, GitLab CI, object storage, Prometheus/Grafana, and third-party APIs.
  • Collaborate with security analysts and the Server team on architecture and production readiness.

Requirements

  • 5+ years of professional backend, platform, or infrastructure engineering experience.
  • Demonstrated experience building and operating multi-stage data or automation pipelines in production.
  • Deep experience with at least one of Python, Go, or Rust.
  • Strong systems design judgment and experience designing for failure and correctness.
  • Production experience with workflow orchestration or job scheduling tools such as Airflow, Temporal, Prefect, Dagster, Argo, or custom schedulers.
  • Practical reliability engineering knowledge, including idempotency, retries, checkpointing, resumability, backpressure, and partial-failure handling.
  • Hands-on observability experience with Prometheus/Grafana, LGTM, or equivalent, including metric design.
  • Deep CI/CD experience, preferably GitLab CI, plus Docker and container-based test environments.
  • Experience with S3/Ceph or equivalent object storage and ClickHouse or another columnar database.
  • Strong debugging, communication, distributed-team, and written and spoken English skills.
  • Experience with progressive delivery, production AI/LLM systems, fleet-scale telemetry, WordPress/PHP/WAF concepts, or configuration management is preferred.

Benefits

  • Fully remote work with flexible hours from anywhere worldwide.
  • 24 paid vacation days, 10 national holidays, and unlimited sick leave.
  • Compensation for private medical insurance.
  • Co-working and gym or sports reimbursement.
  • Professional development and education budget.
  • Opportunity to receive a reward for an innovative patentable idea.
  • Work on challenging engineering projects.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

AI Platform Architect – Anthropic - UK

NewRocket 251-1K Internet Software & Services

NewRocket is seeking an AI Platform Engineer for its AI Foundry to build and operate secure, scalable infrastructure that moves Claude-powered and enterprise AI solutions from prototypes into governed production use.

Argo CD AWS Azure Bash CI/CD CloudFormation Databricks Docker Elasticsearch GCP Generative AI Git GitHub Actions Go GraphQL Helm Java JavaScript Jenkins Kafka Kubeflow Kubernetes Machine Learning Microservices MLflow MLOps MongoDB Network Security OpenSearch PostgreSQL Pulumi Python REST API Secrets Management Serverless Snowflake Terraform TypeScript Vertex AI
2 hours, 39 minutes ago

Staff Platform Engineer

Catena Clearing Internet Software & Services

Catena is hiring a Staff-level infrastructure engineering leader to own the cloud infrastructure, distributed event pipelines, and developer tooling behind its real-time universal fleet telematics data platform.

AWS AWS CDK Docker Kafka Microservices OAuth PostgreSQL Pulumi Python Redis REST API Terraform
3 days, 2 hours ago

Sr. Power Platform Developer

TrueTandem 51-250 Internet Software & Services

TrueTandem is seeking a Senior Power Platform Developer to modernize mission-critical CDC government systems by delivering secure, scalable, and maintainable Microsoft Power Platform solutions.

Agile Azure C# DevSecOps JavaScript JIRA Microsoft Dynamics 365 .NET REST API Scrum SQL Server
4 days, 2 hours ago

Senior Platform Engineer

Flexential 251-1K Internet Software & Services

Flexential is seeking a Senior Platform Engineer to build and operate resilient, secure, AI-enabled IT platforms across observability, DevOps, ITSM, and integrations, with ownership of critical platform architecture and delivery.

Ansible Argo CD AWS CI/CD Containerd Docker Flux GCP GitLab GitOps Grafana HashiCorp Vault Helm JSON Kubernetes OpenTelemetry Prometheus Python REST API TCP/IP Terraform Zabbix
4 days, 2 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers