Senior Engineering Manager, Site Reliability

2 months ago
Full-time
Lead
Software Development
Upstart

Upstart

Upstart Powered Loans: Personal, Car Refinance & Consolidation Through Upstart, apply online for a fast personal loan, auto refinancing, or debt consolidation. Try our quick rate check today with no impact to your credit! Founded by ex Googlers, Upstar...

Banks
1K-5K
Founded 2012

Description

  • Manage and develop a team focused on incident management, observability, operational readiness, and reliability engineering.
  • Define the SRE charter, priorities, roadmap, and measurable outcomes.
  • Translate reliability strategy into capacity-aware plans with clear ownership, milestones, and success measures.
  • Maintain visibility into delivery health, operational risks, and team performance, and intervene early when needed.
  • Build a resilient operating model through cross-training, delegation, and primary/secondary ownership.
  • Lead improvements to incident detection, response, coordination, communication, recovery, and postmortems.
  • Drive consistent observability practices across metrics, logs, traces, alerting, and service health.
  • Establish scalable operational readiness standards for new services, major launches, and architectural changes.
  • Partner with product, platform, infrastructure, and security teams to embed reliability into engineering workflows.
  • Identify systemic reliability risks and turn incident learnings into durable engineering improvements.

Requirements

  • 5+ years of reliability engineering management experience and 7+ years of experience in software engineering, SRE, infrastructure, or platform engineering.
  • Significant hands-on experience in Site Reliability Engineering, Production Engineering, or a similar production operations role.
  • Direct experience managing an SRE or equivalent reliability function, including strategy, roadmap, operating model, and outcomes.
  • Strong technical depth in distributed systems, cloud infrastructure, observability, and production operations.
  • Experience leading high-severity incident response and improving incident management practices at scale.
  • Demonstrated ability to translate strategy into focused, capacity-aware plans and measurable outcomes.
  • Track record of hiring, developing, and retaining high-performing engineers and engineering leaders.
  • Strong cross-functional leadership and communication skills.
  • Experience operating large-scale, highly available distributed systems (preferred).
  • Experience with Datadog, Grafana, Prometheus, OpenTelemetry, Kubernetes, AWS, or similar tools and cloud-native architectures (preferred).
  • Experience implementing service-level objectives and error-budget practices (preferred).
  • Experience building incident management, operational readiness, or resilience programs across a large engineering organization (preferred).

Benefits

  • Base salary range of $195,300 to $270,400 USD for the U.S. remote role.
  • Target bonuses and annual equity grants that vest quarterly.
  • 401(k) or Group Retirement Savings Plan with company match of $2 for every $1 contributed, up to $15,000 annually.
  • Employee Stock Purchase Plan (US only).
  • Comprehensive medical, dental, and vision coverage, plus wellness resources.
  • Paid time off, sick leave, and company holidays.
  • Paid family and parental leave.
  • Employee Assistance Program with mental health support and life resources.
  • Annual wellness allowance and annual productivity allowance.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Engineering Manager, Forward Deployed Engineering

Warner Music Group is hiring an Engineering Manager for its Forward Deployed Engineering team to lead embedded engineers and build a scalable, governed model for delivering AI and automation solutions across global business units.

Agile Databricks Machine Learning
1 hour, 6 minutes ago

Mission Management, Software Engineering Manager

BlackSky 251-1K Professional Services

BlackSky is seeking a Software Engineering Manager to lead the Mission Management team developing and operating autonomous software that plans imaging and communications across its satellite constellation.

Ansible AWS Azure C++ CI/CD Git Go Kubernetes Microservices Nomad Python Terraform
1 hour, 21 minutes ago

Senior Engineering Manager - SaaS (4045)

GBG 1K-5K Professional Services

GBG is seeking a Senior Engineering Manager to lead three full-stack squads building and operating the Go identity-verification platform, while improving delivery, platform coherence, and AI-enabled engineering practices.

CI/CD
1 hour, 36 minutes ago

Senior Manager, Engineering (Container Product Engineering)

Chainguard 51-250 Internet Software & Services

Chainguard is hiring a Senior Engineering Manager to lead the Container Product Engineering team building secure container images, registry services, build systems, and adoption tools for engineers and AI agents.

AWS Docker GCP Go Kubernetes Linux
1 day, 1 hour ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers