Staff Platform Reliability Engineer

5 months, 1 week ago
Full-time
Lead
DevOps and Infrastructure
Puck

Puck

Puck helps great teams find great teammates through employer branding, conversations, and authentic candidate engagement, using personalized automation to enhance the candidate experience and improve hiring metrics.

Internet Software & Services
1-10
Founded 2020

Description

  • Serve as the technical owner of Tempest, Domino's scale and reliability platform.
  • Diagnose and resolve performance bottlenecks and resource misconfigurations surfaced by scale testing.
  • Profile services and trace root causes using observability data from Prometheus and New Relic.
  • Partner with platform and infrastructure teams to ship durable fixes rather than only filing tickets.
  • Deliver accurate, data-driven sizing recommendations for customer-facing documentation.
  • Strengthen observability by improving instrumentation, dashboards, and queries for scale testing.
  • Establish and operationalize scale testing on cloud platforms with appropriate sizing and configuration guidance.
  • Enable scale and reliability testing across additional cloud providers in partnership with platform teams.
  • Build infrastructure automation that improves operational efficiency as the product and customer base grow.

Requirements

  • Background in SRE, platform engineering, or infrastructure.
  • Hands-on experience operating and troubleshooting distributed systems in production Kubernetes environments.
  • Strong proficiency in Python.
  • Comfort working in a large, modular codebase spanning orchestration, infrastructure automation, and systems integration.
  • Experience with observability stacks such as Prometheus, Grafana, New Relic, or similar.
  • Ability to write queries, build dashboards, and use metrics to diagnose performance and reliability issues.
  • Demonstrated ability to profile services, identify resource bottlenecks, and drive durable fixes with engineering teams.
  • Familiarity with performance and load testing tools or methodologies such as Locust, k6, or similar.
  • Self-directed, accountable ownership mindset.
  • Ability to communicate priorities and status effectively in a remote, async environment.

Benefits

  • Annual US base salary range of $185,000 to $230,000.
  • Additional equity may be included.
  • Company bonus or sales commissions/bonuses may be included.
  • 401(k) plan.
  • Medical, dental, and vision benefits.
  • Wellness stipends.
  • Remote role (#LI-Remote).

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Cloud Site Reliability Engineer

Cadwell 51-250 Health Care Providers & Services

Cadwell is seeking a Cloud Site Reliability Engineer to operate and improve AWS infrastructure supporting healthcare customers and ensure reliable, secure, and compliant hosted neurodiagnostic software environments.

AWS Bash CI/CD Encryption HIPAA JavaScript JSON Python Terraform TypeScript YAML
27 minutes ago

Site Reliability Engineer

GiveCampus 51-250 Internet Software & Services

GiveCampus is seeking a hands-on Site Reliability Engineer to strengthen the reliability, performance, observability, and operational maturity of its AWS-based fundraising platform in a remote-first U.S. role.

AWS CI/CD CircleCI Datadog GitHub Actions Kubernetes Linux New Relic OpenSearch PostgreSQL Redis Ruby Ruby on Rails Terraform
57 minutes ago

DevOps / SRE / DevSecOps Engineer (AWS) - Latin America, Remote

Bluelight Consulting 11-50 Internet Software & Services

Bluelight is hiring a DevOps/SRE/DevSecOps professional to operate secure, reliable AWS-based systems and delivery infrastructure for client software projects in a remote consultancy environment.

AWS CI/CD DevSecOps Docker K6 OpenTelemetry PostgreSQL Secrets Management Terraform
1 day ago

DevOps / SRE / DevSecOps Engineer (AWS) - Latin America, Remote

Bluelight Consulting 11-50 Internet Software & Services

Bluelight is hiring a DevOps/SRE/DevSecOps professional to support complex client systems by building secure, reliable, and observable AWS infrastructure and delivery operations.

AWS CI/CD DevSecOps Docker K6 OpenTelemetry PostgreSQL Secrets Management Terraform
1 day ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers