Senior Engineering Manager, Site Reliability

3 weeks, 3 days ago
Full-time
Lead
Software Development
Upstart

Upstart

Upstart Powered Loans: Personal, Car Refinance & Consolidation Through Upstart, apply online for a fast personal loan, auto refinancing, or debt consolidation. Try our quick rate check today with no impact to your credit! Founded by ex Googlers, Upstar...

Banks
1K-5K
Founded 2012

Description

  • Manage and develop a team focused on incident management, observability, operational readiness, and reliability engineering.
  • Define the SRE charter, priorities, roadmap, and measurable outcomes.
  • Translate reliability strategy into capacity-aware plans with clear ownership, milestones, and success measures.
  • Maintain visibility into delivery health, operational risks, and team performance, and intervene early when needed.
  • Build a resilient operating model through cross-training, delegation, and primary/secondary ownership.
  • Lead improvements to incident detection, response, coordination, communication, recovery, and postmortems.
  • Drive consistent observability practices across metrics, logs, traces, alerting, and service health.
  • Establish scalable operational readiness standards for new services, major launches, and architectural changes.
  • Partner with product, platform, infrastructure, and security teams to embed reliability into engineering workflows.
  • Identify systemic reliability risks and turn incident learnings into durable engineering improvements.

Requirements

  • 5+ years of reliability engineering management experience and 7+ years of experience in software engineering, SRE, infrastructure, or platform engineering.
  • Significant hands-on experience in Site Reliability Engineering, Production Engineering, or a similar production operations role.
  • Direct experience managing an SRE or equivalent reliability function, including strategy, roadmap, operating model, and outcomes.
  • Strong technical depth in distributed systems, cloud infrastructure, observability, and production operations.
  • Experience leading high-severity incident response and improving incident management practices at scale.
  • Demonstrated ability to translate strategy into focused, capacity-aware plans and measurable outcomes.
  • Track record of hiring, developing, and retaining high-performing engineers and engineering leaders.
  • Strong cross-functional leadership and communication skills.
  • Experience operating large-scale, highly available distributed systems (preferred).
  • Experience with Datadog, Grafana, Prometheus, OpenTelemetry, Kubernetes, AWS, or similar tools and cloud-native architectures (preferred).
  • Experience implementing service-level objectives and error-budget practices (preferred).
  • Experience building incident management, operational readiness, or resilience programs across a large engineering organization (preferred).

Benefits

  • Base salary range of $195,300 to $270,400 USD for the U.S. remote role.
  • Target bonuses and annual equity grants that vest quarterly.
  • 401(k) or Group Retirement Savings Plan with company match of $2 for every $1 contributed, up to $15,000 annually.
  • Employee Stock Purchase Plan (US only).
  • Comprehensive medical, dental, and vision coverage, plus wellness resources.
  • Paid time off, sick leave, and company holidays.
  • Paid family and parental leave.
  • Employee Assistance Program with mental health support and life resources.
  • Annual wellness allowance and annual productivity allowance.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Engineering Manager, AI Platform

Vannevar Labs 11-50 Aerospace & Defense

Vannevar is hiring a Senior Engineering Manager to lead its AI Platform organization, building the shared foundation for machine learning and agent systems that power national security workflows.

AWS LLM Machine Learning PostgreSQL Python
20 hours, 57 minutes ago

Manager, Software Engineering - Data Platform

Figma 1K-5K Internet Software & Services

Figma is hiring a Data Platform engineering leader to scale the systems that power core business metrics, product analytics, and applied AI and experimentation across the company.

Machine Learning
20 hours, 57 minutes ago

Engineering Manager - Remote

Formativ Group Internet Software & Services

FormativGroup is hiring a hands-on Engineering Manager to lead a financial services technology team building and modernizing internal products for tax professionals while staying directly involved in delivery.

Angular AWS AWS CDK Azure C# CI/CD GCP GitHub GitHub Actions .NET Node.js Pulumi Python React SonarQube Terraform
20 hours, 57 minutes ago

Engineering Manager – Unified Telephony Platform

Eltropy 51-250 Communications Equipment

Remote Engineering Manager role at a communications platform company leading the buildout of a cloud-native unified telephony platform for financial institutions.

CI/CD Go Microservices Python Twilio
1 day, 21 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers