Lead Site Reliability Engineer (Performance & Scalability) | Contract | Remote US

14 hours, 8 minutes ago
Contract
Lead
DevOps and Infrastructure
Tech Holding

Tech Holding

Tech Holding: California's #1 website design company offering full-service technology consulting with expertise in software management, AI, and security.

Internet Software & Services
51-250
Founded 2016

Description

  • Establish performance, throughput, latency, and capacity baselines for critical platform workflows.
  • Define and maintain SLOs, error budgets, performance budgets, dashboards, alerts, and reliability thresholds.
  • Instrument and analyze the full request path across services, infrastructure, databases, networking, caches, queues, DNS, and third-party dependencies.
  • Identify system bottlenecks and lead cross-functional remediation efforts with engineering teams.
  • Build capacity models that show current limits, emerging constraints, and the cost of additional scale.
  • Lead load, stress, soak, spike, failure, and recovery testing in representative environments.
  • Develop realistic demand scenarios for major customers, partnerships, pilots, and high-volume events.
  • Drive architecture hardening, graceful degradation, dependency-failure planning, and resilience improvements.
  • Partner with Test Automation and Scalability Engineering on automated performance testing, regression coverage, and production release gates.
  • Own technical readiness assessments for major pilots, partnerships, and production launches.
  • Create operational runbooks for scale-up events, incidents, rollback, recovery, and dependency failures.
  • Lead performance and reliability investigations during incidents and incorporate lessons into future engineering work.
  • Communicate infrastructure cost, performance, and reliability tradeoffs to engineering and executive leadership.
  • Recommend capacity and reliability investments before they become production constraints.

Requirements

  • Significant experience in Site Reliability Engineering, performance engineering, platform engineering, distributed systems, or a closely related discipline.
  • Experience supporting production systems with meaningful scale, traffic, latency, or availability requirements.
  • Deep understanding of observability, performance analysis, capacity planning, and reliability engineering.
  • Strong hands-on experience with cloud infrastructure and production distributed systems.
  • Deep knowledge of databases, networking, caching, queueing, compute, storage, and common distributed-system failure modes.
  • Experience defining and operating against SLOs, SLIs, error budgets, and production reliability metrics.
  • Hands-on experience performing load, stress, soak, scalability, and resilience testing.
  • Ability to profile systems, diagnose bottlenecks, tune architecture, and work directly with engineering teams to implement improvements.
  • Experience designing for graceful degradation, dependency failures, recovery, and high-demand scenarios.
  • Strong incident management and root-cause analysis experience.
  • Ability to translate technical performance and reliability risks into clear business implications for senior leadership.
  • Strong judgment around when systems need optimization versus when added complexity is premature.
  • Experience operating high-scale SaaS, identity, DNS, registry, infrastructure, or other highly distributed platforms (preferred).
  • Experience creating capacity-cost models and forecasting infrastructure requirements (preferred).
  • Experience building performance and reliability gates into CI/CD pipelines (preferred).
  • Experience preparing platforms for major traffic increases from enterprise customers or strategic partnerships (preferred).
  • Experience leading reliability or performance initiatives across multiple engineering teams (preferred).
  • Applicants must be authorized to work for any employer in the U.S.; no visa sponsorship is available.

Benefits

  • Remote work opportunity (#LI-Remote).
  • Contract employment type.
  • Equal Opportunity Employer commitment.
  • Inclusive workplace across all backgrounds and experiences.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Site Reliability Engineer

Coforge 10K-50K IT Services

Coforge is hiring a remote Site Reliability Engineer to help build and operate reliable cloud platforms and production systems for teams across Costa Rica, Peru, Colombia, and Bolivia.

AWS Azure C# CI/CD Datadog GCP Go Grafana Java Kubernetes Linux Microservices New Relic OpenTelemetry Prometheus Python Splunk Terraform TypeScript
1 day, 14 hours ago

AMS:US:SL:Service Reliability Engineer:Lead

Thoughtworks 10K-50K Professional Services

Thoughtworks is hiring a Service Reliability Engineer to improve the reliability, resilience, and performance of client infrastructure and production systems.

Ansible CircleCI CloudFormation ELK Stack GitLab GitOps Go Grafana Jaeger Java Jenkins Kubernetes Nomad Prometheus Python Ruby Shell Scripting Terraform Zipkin
1 day, 14 hours ago

Sr. Staff Platform/Data Reliability Engineer, Databricks (R5537)

Bitly 51-250 Internet Software & Services

Shield AI is seeking a Sr. Staff Platform / Data Reliability Engineer to make its Databricks platform reliable, secure, scalable, and operationally mature for enterprise and regulated use.

CI/CD Databricks Git
3 days, 14 hours ago

Senior Manager Site Reliability Engineer

GoDaddy 5001-10000 Technology, Information and Internet

GoDaddy is hiring a Senior Manager of Site Reliability Engineering in India to lead its SRE Center of Excellence and build a global, platform-focused operations capability that improves reliability, developer productivity, and scale.

AWS CI/CD Datadog GitHub Actions Go Grafana Kafka Kubernetes Microservices PostgreSQL Prometheus Python Terraform
4 days, 14 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers