Database Reliability Engineer - Core Team

3 months, 4 weeks ago
ClickHouse

ClickHouse

ClickHouse provides a fast open source column-oriented database management system that enables users to generate real-time analytical data reports through SQL queries, catering to the needs of industries requiring efficient data processing and analysis.

IT Services
51-250
Founded 2021
$300M raised

Description

  • Build and lead processes to improve the reliability, availability, scalability, and performance of ClickHouse Core.
  • Collaborate with Control Plane, Dataplane, Security, Support, and Operations teams to implement ClickHouse effectively for customers.
  • Own engineering escalation management, incident response, investigations, and response coordination.
  • Conduct post-mortem analysis, including running blameless postmortems, and drive continuous improvement.
  • Improve metrics and alerts to detect and prevent production issues before they impact customers.
  • Investigate common customer problems, identify root causes, and submit bug fixes, issue reports, and improvement suggestions.
  • Enhance incident response processes for core-related outages and communicate with impacted customers alongside Support and Cloud teams.
  • Plan, enable, and drive chaos engineering initiatives across engineering teams.
  • Manage on-call processes for performance and reliability issues and establish escalation best practices.

Requirements

  • Bachelor’s or Master’s degree in Computer Science or a related field.
  • At least 5 years of experience in Reliability Engineering, QA, or customer-facing engineering.
  • Experience operating ClickHouse or other SQL databases in production.
  • Strong understanding of distributed database internals and SQL, with ClickHouse experience being a major plus.
  • Scripting experience with Shell or Python.
  • Ability to read and understand C++ code.
  • Knowledge of cloud platforms such as AWS, Azure, or Google Cloud Platform.
  • Strong problem-solving and production debugging skills.
  • Experience working effectively in a fast-paced global team with high ownership and accountability.
  • Excellent communication skills.

Benefits

  • Remote-friendly flexible work environment across multiple countries, including the Netherlands, UK, United States, and Germany.
  • Employer contributions toward healthcare.
  • Stock options for every new team member.
  • Flexible time off in the US and generous time off in other countries.
  • A $500 home office setup allowance for remote employees.
  • Opportunities to attend company-wide global gatherings and offsites.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Manager Site Reliability Engineer

GoDaddy 5001-10000 Technology, Information and Internet

GoDaddy is hiring a Senior Manager of Site Reliability Engineering in India to lead its SRE Center of Excellence and build a global, platform-focused operations capability that improves reliability, developer productivity, and scale.

AWS CI/CD Datadog GitHub Actions Go Grafana Kafka Kubernetes Microservices PostgreSQL Prometheus Python Terraform
7 hours, 35 minutes ago

Vice President, Global Production Operations & Reliability

Everbridge 1K-5K Internet Software & Services

Everbridge is hiring a Vice President, Global Production Operations & Reliability to lead the company’s global production operations for its cloud-native SaaS platform and drive reliability, scalability, security, and operational excellence.

AWS CI/CD Kubernetes
1 day, 6 hours ago

DevOps Engineer - SRE Observability

Lingaro 5K-10K IT Services

An infrastructure-focused role at Lingaro responsible for monitoring, automating, and designing cloud systems within an Azure-based environment.

Azure Azure Pipelines CI/CD Docker GitHub GitHub Actions Grafana Kubernetes MySQL PostgreSQL Prometheus SQL Terraform
1 day, 6 hours ago

Staff Field Reliability Engineer

Honeycomb.io 51-250 Internet Software & Services

Honeycomb is hiring a Field Reliability Engineer to lead complex customer escalations, managed infrastructure operations, and observability strategy for its cloud-based platform.

Ansible AWS Chef EC2 Go Helm Honeycomb Java Kubernetes .NET Node.js OpenTelemetry Python Serverless Terraform TypeScript
1 day, 7 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers