CloudLinux

CloudLinux

CloudLinux is a leading provider of the CloudLinux OS, a platform for Linux web hosting that offers next-level performance and security. With a focus on optimizing web hosting environments, CloudLinux helps service providers improve density, stability,...

IT Services
51-250
Founded 2009

Description

  • Design and implement a self-service DBaaS platform using Terraform and Ansible for deploying highly available PostgreSQL, ClickHouse, MongoDB, and Redis clusters.
  • Build and operate database infrastructure across bare metal, OpenNebula, Kubernetes, and public cloud environments.
  • Manage and scale large ClickHouse analytics clusters, including sharding, replication, table engine optimization, and S3 backup pipelines.
  • Maintain and scale Apache Airflow and Redash infrastructure to support reliable ETL pipelines and analytics workflows.
  • Implement SRE practices for data management, including automated self-healing and defined SLO/SLI for databases.
  • Lead migration from legacy database solutions to modern cloud-native patterns and help evaluate Kubernetes operators for stateful workloads.
  • Serve as a technical authority for product teams on data schema design and SQL query optimization for high-load systems.
  • Collaborate with infrastructure and analytics teams to improve reliability, observability, and performance across the data platform.
  • Automate infrastructure and operational tasks with code to reduce manual intervention and repeat work.

Requirements

  • 5+ years of deep PostgreSQL experience, including MVCC internals, locking mechanics, Patroni, PgBouncer, and major version upgrades under load.
  • Proven experience operating large ClickHouse clusters, including ZooKeeper or ClickHouse Keeper, sharding, replication internals, and performance troubleshooting.
  • Strong Terraform and Ansible experience, including writing complex modules and roles.
  • Programming experience in Python or Go for infrastructure and automation is a major plus.
  • Experience working in hybrid environments across bare metal, Kubernetes, and cloud platforms.
  • Understanding of database performance tuning, including NVMe and network storage optimization.
  • Systems-level thinking across networking, infrastructure, and application logic.
  • Knowledge of security and disaster recovery practices, including FIPS and audit logs.
  • Preferred experience building an Internal Developer Platform (IDP).
  • Preferred experience operating databases in Kubernetes using CloudNativePG or Altinity Operator.
  • Preferred experience working for cloud or hosting providers in similar service environments.

Benefits

  • Fully remote work with flexible working hours and the ability to work from anywhere worldwide.
  • Paid 24 days of vacation per year.
  • 10 days of national holidays.
  • Unlimited sick leave.
  • Private medical insurance coverage.
  • Co-working and gym/sports reimbursement.
  • Budget for education, training, and conferences.
  • Opportunity to be rewarded for innovative ideas that the company can patent.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

PENN Entertainment 10K-50K Hotels, Restaurants & Leisure

PENN Entertainment is seeking a Senior Site Reliability Engineer to build and operate the cloud infrastructure supporting large-scale sports betting and media platforms across regulated jurisdictions.

Argo CD AWS Bash CI/CD Datadog GCP GitHub Actions GitOps Go Helm Kubernetes Linux PostgreSQL Python Shell Scripting Terraform
17 hours, 6 minutes ago

Staff Site Reliability Engineer, Ads

Reddit 1K-5K Internet Software & Services

Reddit is hiring a Staff Site Reliability Engineer to provide technical leadership for reliability, scalability, and operational excellence across its advertising infrastructure and revenue-critical systems.

Apache Spark ClickHouse GCP Go Kafka Kubernetes
17 hours, 36 minutes ago

Sr Lead Network Reliability Engineer

Coupa Software 1K-5K Internet Software & Services

Coupa is hiring a Sr. Lead Network Development Engineer to scale and operate its global SaaS platform’s cloud networking infrastructure through automation, reliability engineering, and technical leadership.

Ansible AWS Azure Chef DNS Fortinet Go Java Kubernetes Linux Python Ruby TCP/IP Terraform TLS
1 day, 17 hours ago

Senior Monitoring/Observability Architect

Makpar 51-250 Internet Software & Services

Makpar is seeking a Senior Monitoring/Observability Architect to lead enterprise monitoring strategy, architecture, and implementation guidance for a large federal government program.

Datadog Splunk
1 day, 17 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers