Database Reliability Engineer - Core Team

4 months, 1 week ago
ClickHouse

ClickHouse

ClickHouse provides a fast open source column-oriented database management system that enables users to generate real-time analytical data reports through SQL queries, catering to the needs of industries requiring efficient data processing and analysis.

IT Services
51-250
Founded 2021
$300M raised

Description

  • Build and lead processes to improve the reliability, availability, scalability, and performance of ClickHouse Core.
  • Collaborate with Control Plane, Dataplane, Security, Support, and Operations teams to implement ClickHouse effectively for customers.
  • Own engineering escalation management, incident response, investigations, and response coordination.
  • Conduct post-mortem analysis, including running blameless postmortems, and drive continuous improvement.
  • Improve metrics and alerts to detect and prevent production issues before they impact customers.
  • Investigate common customer problems, identify root causes, and submit bug fixes, issue reports, and improvement suggestions.
  • Enhance incident response processes for core-related outages and communicate with impacted customers alongside Support and Cloud teams.
  • Plan, enable, and drive chaos engineering initiatives across engineering teams.
  • Manage on-call processes for performance and reliability issues and establish escalation best practices.

Requirements

  • Bachelor’s or Master’s degree in Computer Science or a related field.
  • At least 5 years of experience in Reliability Engineering, QA, or customer-facing engineering.
  • Experience operating ClickHouse or other SQL databases in production.
  • Strong understanding of distributed database internals and SQL, with ClickHouse experience being a major plus.
  • Scripting experience with Shell or Python.
  • Ability to read and understand C++ code.
  • Knowledge of cloud platforms such as AWS, Azure, or Google Cloud Platform.
  • Strong problem-solving and production debugging skills.
  • Experience working effectively in a fast-paced global team with high ownership and accountability.
  • Excellent communication skills.

Benefits

  • Remote-friendly flexible work environment across multiple countries, including the Netherlands, UK, United States, and Germany.
  • Employer contributions toward healthcare.
  • Stock options for every new team member.
  • Flexible time off in the US and generous time off in other countries.
  • A $500 home office setup allowance for remote employees.
  • Opportunities to attend company-wide global gatherings and offsites.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Site Reliability Engineer

AssureSoft 51-250 Internet Software & Services

AssureSoft is hiring a remote Site Reliability Engineer to support production cloud infrastructure and platform reliability for long-term client projects.

Argo CD AWS Bash DNS Docker GCP GitHub Actions Go Grafana HTTP Kafka Kubernetes Linux Load Balancing Prometheus Python RabbitMQ Snowflake TCP/IP TLS TypeScript Unix
3 days, 15 hours ago

Senior Site Reliability Engineer

Megaport 251-1K Diversified Telecommunication Services

Megaport is hiring a Senior Platform Engineer to support secure, reliable, and maintainable global production systems within its SRE-focused platform team.

AWS Bash Cassandra CI/CD ClickHouse Git GitHub Go Kubernetes Linux PostgreSQL Python Terraform
3 days, 16 hours ago

Site Reliability Engineer - Azure, Observability and Scripting

Jalasoft 1K-5K Internet Software & Services

Jalasoft is hiring a Site Reliability Engineer to support the reliability, scalability, and performance of cloud-native platforms on Microsoft Azure and Kubernetes.

Azure Bash Grafana Kubernetes OpenTelemetry Oracle PowerShell Prometheus Python Terraform
4 days, 16 hours ago

Senior Site Reliability Engineer

Aspenview Technology Partners Internet Software & Services

AspenView Technology Partners is hiring a Senior Site Reliability Engineer to support a large-scale cloud transformation by building and operating resilient, secure platforms for critical enterprise applications.

Agile AWS Bash Datadog DevSecOps GCP Grafana Kubernetes Prometheus Python Splunk Terraform
5 days, 15 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers