Alex Staff Agency

Alex Staff Agency

Alex Staff Agency is a prominent player in the international IT recruitment market, providing staffing and recruiting services for large technical companies worldwide. With a team of highly experienced professional recruiters from various countries and...

Professional Services
11-50
Founded 2018

Description

  • Own production PostgreSQL reliability, including high availability, replication, failover, upgrades, backups, PITR, and restore validation.
  • Improve disaster recovery with tested restores, documented recovery paths, clear RTO/RPO targets, runbooks, and safe maintenance plans.
  • Support the wider database estate, including ClickHouse, MongoDB, and Redis, through incident troubleshooting and operational improvements.
  • Automate DBA workflows for provisioning, grants, backups, restores, health checks, and ownership metadata using infrastructure and scripting tools.
  • Build self-service database capabilities that reduce manual DBA intervention for requests, access, credentials, and operational checks.
  • Improve observability and incident response with dashboards, metrics, logs, alerting, routing, and clear communication during production issues.
  • Review access and data-safety changes and help maintain safe, auditable operational processes.
  • Reduce single-person dependency by documenting production patterns and sharing operational knowledge across the team.

Requirements

  • 5+ years of deep hands-on PostgreSQL experience in business-critical production environments, or equivalent depth.
  • Strong PostgreSQL operations knowledge, including MVCC, WAL, transactions, locks, indexes, query planning, replication, autovacuum, bloat, major upgrades, backups, PITR, and restore testing.
  • Experience with highly available databases and understanding of quorum, split-brain risk, failover, rollback, and recovery.
  • Strong Linux and infrastructure fundamentals, including systemd, networking, storage, filesystems, CPU/memory/disk bottlenecks, TLS, DNS, firewalls, and root-cause troubleshooting.
  • Automation experience with Ansible and scripting; Terraform/OpenTofu, GitLab CI/CD, and merge-request-based delivery are strong advantages.
  • Ability and willingness to support more than one database engine and learn ClickHouse quickly.
  • Practical use of AI engineering assistants such as Claude and Codex, with careful human verification of generated output.
  • Upper-intermediate English or higher for clear team communication.
  • ClickHouse experience is a strong plus, though not required on day one.

Benefits

  • Fully remote work with flexible working hours worldwide.
  • 24 paid vacation days per year, plus 10 national holidays and unlimited sick leave.
  • Private medical insurance reimbursement.
  • Co-working and gym/sports reimbursement.
  • Budget for education.
  • Focus on professional development.
  • Opportunity to work on interesting and challenging projects.
  • Reward for the most innovative patentable idea.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

GoDaddy 5001-10000 Technology, Information and Internet

GoDaddy is hiring a Senior Site Reliability Engineer to support the Commerce ecosystem remotely by improving the reliability, scalability, security, and operation of business-critical production platforms.

Ansible AWS AWS CDK CI/CD CloudFormation Go Kubernetes Linux Pulumi Python SaltStack Terraform TypeScript
3 hours, 32 minutes ago

Senior Site Reliability Engineer

PENN Entertainment 10K-50K Hotels, Restaurants & Leisure

PENN Entertainment is seeking a Senior Site Reliability Engineer to build and operate the cloud infrastructure supporting large-scale sports betting and media platforms across regulated jurisdictions.

Argo CD AWS Bash CI/CD Datadog GCP GitHub Actions GitOps Go Helm Kubernetes Linux PostgreSQL Python Shell Scripting Terraform
1 day, 3 hours ago

Staff Site Reliability Engineer, Ads

Reddit 1K-5K Internet Software & Services

Reddit is hiring a Staff Site Reliability Engineer to provide technical leadership for reliability, scalability, and operational excellence across its advertising infrastructure and revenue-critical systems.

Apache Spark ClickHouse GCP Go Kafka Kubernetes
1 day, 3 hours ago

Sr Lead Network Reliability Engineer

Coupa Software 1K-5K Internet Software & Services

Coupa is hiring a Sr. Lead Network Development Engineer to scale and operate its global SaaS platform’s cloud networking infrastructure through automation, reliability engineering, and technical leadership.

Ansible AWS Azure Chef DNS Fortinet Go Java Kubernetes Linux Python Ruby TCP/IP Terraform TLS
2 days, 3 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers