Staff Site Reliability Engineer, Ads

4 weeks, 1 day ago
Full-time
Lead
DevOps and Infrastructure
Reddit

Reddit

Reddit is a diverse network of communities where users engage in content creation, voting, and discussions on a wide range of topics.

Internet Software & Services
1K-5K
Founded 2005
$550M raised

Description

  • Lead reliability initiatives across ad serving, auctions, targeting, reporting, measurement, and billing domains.
  • Partner with engineering leadership to develop reliability, scalability, and operational excellence roadmaps.
  • Design and build platforms, tooling, and automation that improve reliability and developer productivity.
  • Lead architecture reviews and influence technical decisions for critical advertising systems.
  • Participate in on-call rotations, investigate complex incidents, and coordinate major production responses.
  • Identify systemic reliability risks and drive long-term resilience improvements.
  • Establish reliability metrics and SLOs for advertiser-critical user journeys.
  • Mentor engineers and provide technical leadership across multiple teams.
  • Influence product and infrastructure planning to incorporate reliability considerations.

Requirements

  • 8+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or related roles operating large-scale distributed systems.
  • Experience evolving high-traffic, user-facing production environments.
  • Deep expertise in distributed systems, scale engineering, and cloud-native architectures.
  • Experience designing highly available systems and implementing strong operational practices.
  • Strong software engineering skills in general-purpose backend languages such as Go.
  • Understanding of observability systems, including metrics, logging, tracing, and alerting.
  • Experience with SLOs, automation, incident management, troubleshooting, and performance optimization.
  • Strong collaboration and communication skills with the ability to influence technical direction.
  • Experience with advertising technology or other revenue-critical systems (preferred).
  • Experience with Kubernetes, cloud infrastructure, Kafka, ClickHouse, Spark, Flink, BigQuery, machine learning systems, or similar technologies (preferred).

Benefits

  • Remote-friendly, flexible workforce.
  • Base salary range of $217,000–$303,900 USD.
  • Equity in the form of restricted stock units; commission may apply depending on the position.
  • Comprehensive medical, dental, and vision benefits.
  • 401(k) with employer matching.
  • Flexible vacation, Reddit Global Days Off, and paid volunteer time.
  • Four or more months of paid parental leave and family planning support.
  • Home-office benefits and personal and professional development funds.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Arquitecto de soluciones / SRE - Argentina

Coderio 51-250 Internet Software & Services

Coderio busca un/a Arquitecto/a de Soluciones / SRE en Argentina para un entorno bancario, responsable de configurar, monitorear y analizar soluciones orientadas a pruebas de performance.

AWS Grafana Java Jenkins Kafka Kubernetes Microservices .NET Node.js OpenShift Prometheus Python RabbitMQ
1 hour, 54 minutes ago

SRE Engineer Contractor

Beauty For All Industries (BFA) 51-200 information technology & services

IPSY is seeking a remote Site Reliability Engineer Contractor in Mexico or Colombia, covering the PST time zone, to improve the availability, resilience, and operational reliability of its beauty membership platform.

Amplitude AWS Bash CI/CD Contentful Datadog Grafana Microservices Netlify New Relic OpsGenie PagerDuty Prometheus Python Terraform
2 days, 1 hour ago

[Job 31894] AI ORCHESTRATOR (APP SRE)

CI&T 5K-10K Internet Software & Services

A CI&T busca um AI Orchestrator para governar a entrega ponta a ponta de um projeto de monitoramento SRE em jornadas de aplicativos de saúde, coordenando equipes e agentes de IA desde o upstream até a produção.

Agile
4 days, 2 hours ago

Senior Site Reliability Engineer (SRE)

UJET 251-1K Professional Services

UJET is seeking a Senior Site Reliability Engineer to build and scale its SRE function, improving reliability, reducing operational toil, and establishing production engineering best practices for its cloud-based contact center platform.

AWS Azure Go Java Kubernetes Microservices Python Terraform
5 days, 1 hour ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers