Alex Staff Agency

Alex Staff Agency

Alex Staff Agency is a prominent player in the international IT recruitment market, providing staffing and recruiting services for large technical companies worldwide. With a team of highly experienced professional recruiters from various countries and...

Professional Services
11-50
Founded 2018

Description

  • Own production PostgreSQL reliability, including high availability, replication, failover, upgrades, backups, PITR, and restore validation.
  • Improve disaster recovery with tested restores, documented recovery paths, clear RTO/RPO targets, runbooks, and safe maintenance plans.
  • Support the wider database estate, including ClickHouse, MongoDB, and Redis, through incident troubleshooting and operational improvements.
  • Automate DBA workflows for provisioning, grants, backups, restores, health checks, and ownership metadata using infrastructure and scripting tools.
  • Build self-service database capabilities that reduce manual DBA intervention for requests, access, credentials, and operational checks.
  • Improve observability and incident response with dashboards, metrics, logs, alerting, routing, and clear communication during production issues.
  • Review access and data-safety changes and help maintain safe, auditable operational processes.
  • Reduce single-person dependency by documenting production patterns and sharing operational knowledge across the team.

Requirements

  • 5+ years of deep hands-on PostgreSQL experience in business-critical production environments, or equivalent depth.
  • Strong PostgreSQL operations knowledge, including MVCC, WAL, transactions, locks, indexes, query planning, replication, autovacuum, bloat, major upgrades, backups, PITR, and restore testing.
  • Experience with highly available databases and understanding of quorum, split-brain risk, failover, rollback, and recovery.
  • Strong Linux and infrastructure fundamentals, including systemd, networking, storage, filesystems, CPU/memory/disk bottlenecks, TLS, DNS, firewalls, and root-cause troubleshooting.
  • Automation experience with Ansible and scripting; Terraform/OpenTofu, GitLab CI/CD, and merge-request-based delivery are strong advantages.
  • Ability and willingness to support more than one database engine and learn ClickHouse quickly.
  • Practical use of AI engineering assistants such as Claude and Codex, with careful human verification of generated output.
  • Upper-intermediate English or higher for clear team communication.
  • ClickHouse experience is a strong plus, though not required on day one.

Benefits

  • Fully remote work with flexible working hours worldwide.
  • 24 paid vacation days per year, plus 10 national holidays and unlimited sick leave.
  • Private medical insurance reimbursement.
  • Co-working and gym/sports reimbursement.
  • Budget for education.
  • Focus on professional development.
  • Opportunity to work on interesting and challenging projects.
  • Reward for the most innovative patentable idea.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Cloud Site Reliability Engineer

Cadwell 51-250 Health Care Providers & Services

Cadwell is seeking a Cloud Site Reliability Engineer to operate and improve AWS infrastructure supporting healthcare customers and ensure reliable, secure, and compliant hosted neurodiagnostic software environments.

AWS Bash CI/CD Encryption HIPAA JavaScript JSON Python Terraform TypeScript YAML
21 hours, 37 minutes ago

Site Reliability Engineer

GiveCampus 51-250 Internet Software & Services

GiveCampus is seeking a hands-on Site Reliability Engineer to strengthen the reliability, performance, observability, and operational maturity of its AWS-based fundraising platform in a remote-first U.S. role.

AWS CI/CD CircleCI Datadog GitHub Actions Kubernetes Linux New Relic OpenSearch PostgreSQL Redis Ruby Ruby on Rails Terraform
22 hours, 7 minutes ago

DevOps / SRE / DevSecOps Engineer (AWS) - Latin America, Remote

Bluelight Consulting 11-50 Internet Software & Services

Bluelight is hiring a DevOps/SRE/DevSecOps professional to operate secure, reliable AWS-based systems and delivery infrastructure for client software projects in a remote consultancy environment.

AWS CI/CD DevSecOps Docker K6 OpenTelemetry PostgreSQL Secrets Management Terraform
1 day, 21 hours ago

DevOps / SRE / DevSecOps Engineer (AWS) - Latin America, Remote

Bluelight Consulting 11-50 Internet Software & Services

Bluelight is hiring a DevOps/SRE/DevSecOps professional to support complex client systems by building secure, reliable, and observable AWS infrastructure and delivery operations.

AWS CI/CD DevSecOps Docker K6 OpenTelemetry PostgreSQL Secrets Management Terraform
1 day, 21 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers