Senior Site Reliability Engineer (Performance and Scalability)

14 hours, 14 minutes ago
Full-time
Senior
DevOps and Infrastructure
Digitalzone

Digitalzone

Digitalzone is a leading B2B marketing agency specializing in personalized and data-driven campaigns for rapid business growth.

Media
251-1K
Founded 2013

Description

  • Build capacity planning, autoscaling, caching, queueing, and graceful-degradation capabilities for campaign spikes.
  • Establish load and failure testing practices, including frameworks, tooling, and runbooks for engineering teams.
  • Own SLOs, error budgets, and observability across TypeScript, Go, and PHP/Laravel services.
  • Standardize instrumentation for scale across engineering teams.
  • Improve PostgreSQL and AWS performance and availability through automation and infrastructure as code.
  • Lead incident response and blameless postmortems, driving systemic reliability improvements.
  • Partner with teams early on capacity and resilience planning to help them scale services independently.

Requirements

  • 5+ years of experience in SRE, platform, or backend engineering.
  • Production ownership of large-scale systems handling tens of thousands of requests per minute.
  • Experience scaling systems through real traffic spikes.
  • Experience designing and running load and failure testing programs adopted by other teams.
  • Deep AWS experience and strong knowledge of PostgreSQL performance and scaling.
  • Fluency with observability tools and infrastructure as code.
  • Scripting experience with Go, TypeScript, or similar languages.
  • Calm, systematic incident-management approach and strong communication skills for cross-team enablement.

Benefits

  • Immediate, large-scale impact on a high-growth business.
  • Top-of-market compensation package.
  • Opportunity to work with experienced talent from Talabat, Careem, Etisalat, and other regional companies.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Incident management / reliability / SRE Evaluator

Weekday 11-50 Construction & Engineering

An independent contractor Evaluator will remotely assess AI-generated documents, spreadsheets, and presentations for accuracy, rigor, and quality using incident management, reliability, and SRE expertise.

14 hours, 29 minutes ago

DevOps / SRE / DevSecOps Engineer (AWS) - Remote, Latin America

Bluelight Consulting 11-50 Internet Software & Services

Bluelight is hiring a DevOps/SRE/DevSecOps professional to build, operate, and secure reliable cloud applications and infrastructure for client projects at its growing software consultancy.

AWS CI/CD DevSecOps Docker K6 OpenTelemetry PostgreSQL Secrets Management Terraform
1 day, 13 hours ago

DevOps / SRE / DevSecOps Engineer (AWS) - Remote, Latin America

Bluelight Consulting 11-50 Internet Software & Services

Bluelight is seeking a DevOps/SRE professional to support client projects by building, securing, and operating reliable AWS-based applications and infrastructure.

AWS CI/CD DevSecOps Docker K6 OpenTelemetry PostgreSQL Secrets Management Terraform
1 day, 13 hours ago

DevOps / SRE / DevSecOps Engineer (AWS) - Remote, Latin America

Bluelight Consulting 11-50 Internet Software & Services

Bluelight is seeking a DevOps, SRE, or DevSecOps professional to support client projects by building secure, reliable, and scalable cloud infrastructure and delivery systems.

AWS CI/CD DevSecOps Docker K6 OpenTelemetry PostgreSQL Secrets Management Terraform
1 day, 13 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers