Staff Site Reliability Engineer, Ads

16 hours, 43 minutes ago
Full-time
Lead
DevOps and Infrastructure
Reddit

Reddit

Reddit is a diverse network of communities where users engage in content creation, voting, and discussions on a wide range of topics.

Internet Software & Services
1K-5K
Founded 2005
$550M raised

Description

  • Lead reliability initiatives across ad serving, auctions, targeting, reporting, measurement, and billing domains.
  • Partner with engineering leadership to develop reliability, scalability, and operational excellence roadmaps.
  • Design and build platforms, tooling, and automation that improve reliability and developer productivity.
  • Lead architecture reviews and influence technical decisions for critical advertising systems.
  • Participate in on-call rotations, investigate complex incidents, and coordinate major production responses.
  • Identify systemic reliability risks and drive long-term resilience improvements.
  • Establish reliability metrics and SLOs for advertiser-critical user journeys.
  • Mentor engineers and provide technical leadership across multiple teams.
  • Influence product and infrastructure planning to incorporate reliability considerations.

Requirements

  • 8+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or related roles operating large-scale distributed systems.
  • Experience evolving high-traffic, user-facing production environments.
  • Deep expertise in distributed systems, scale engineering, and cloud-native architectures.
  • Experience designing highly available systems and implementing strong operational practices.
  • Strong software engineering skills in general-purpose backend languages such as Go.
  • Understanding of observability systems, including metrics, logging, tracing, and alerting.
  • Experience with SLOs, automation, incident management, troubleshooting, and performance optimization.
  • Strong collaboration and communication skills with the ability to influence technical direction.
  • Experience with advertising technology or other revenue-critical systems (preferred).
  • Experience with Kubernetes, cloud infrastructure, Kafka, ClickHouse, Spark, Flink, BigQuery, machine learning systems, or similar technologies (preferred).

Benefits

  • Remote-friendly, flexible workforce.
  • Base salary range of $217,000–$303,900 USD.
  • Equity in the form of restricted stock units; commission may apply depending on the position.
  • Comprehensive medical, dental, and vision benefits.
  • 401(k) with employer matching.
  • Flexible vacation, Reddit Global Days Off, and paid volunteer time.
  • Four or more months of paid parental leave and family planning support.
  • Home-office benefits and personal and professional development funds.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

PENN Entertainment 10K-50K Hotels, Restaurants & Leisure

PENN Entertainment is seeking a Senior Site Reliability Engineer to build and operate the cloud infrastructure supporting large-scale sports betting and media platforms across regulated jurisdictions.

Argo CD AWS Bash CI/CD Datadog GCP GitHub Actions GitOps Go Helm Kubernetes Linux PostgreSQL Python Shell Scripting Terraform
16 hours, 13 minutes ago

Sr Lead Network Reliability Engineer

Coupa Software 1K-5K Internet Software & Services

Coupa is hiring a Sr. Lead Network Development Engineer to scale and operate its global SaaS platform’s cloud networking infrastructure through automation, reliability engineering, and technical leadership.

Ansible AWS Azure Chef DNS Fortinet Go Java Kubernetes Linux Python Ruby TCP/IP Terraform TLS
1 day, 16 hours ago

Senior Monitoring/Observability Architect

Makpar 51-250 Internet Software & Services

Makpar is seeking a Senior Monitoring/Observability Architect to lead enterprise monitoring strategy, architecture, and implementation guidance for a large federal government program.

Datadog Splunk
1 day, 16 hours ago

Staff Site Reliability Engineer

Caseware 251-1K Internet Software & Services

Caseware is hiring a Senior Site Reliability/Platform Engineer for its Canada-remote team to improve production resilience, security, operational excellence, and developer enablement across its audit and accounting software platform.

AWS AWS CDK CI/CD GitHub Actions Kubernetes Node.js OpenTelemetry TypeScript
1 day, 16 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers