Site Reliability Engineer

5 days, 17 hours ago
Full-time
Senior
DevOps and Infrastructure

GoDaddy

GoDaddy is a technology company in the Technology, Information and Internet industry. It describes itself as the world’s largest web services platform and a provider of tools for domains, websites, marketing, email, security, and online business services.

Technology, Information and Internet
5001-10000
Founded 1997

Description

  • Operate, scale, troubleshoot, and improve OpenStack-based compute, networking, and storage services in large-scale production environments.
  • Drive OpenStack migration efforts by designing migration tooling, validation processes, and rollback strategies.
  • Build and maintain Python and Puppet automation to eliminate manual operational work.
  • Improve monitoring, alerting, dashboards, and overall system observability.
  • Participate in shared on-call rotations and lead incident response and blameless post-incident reviews.
  • Review code and technical designs, document systems and runbooks, and mentor SRE I/II engineers.
  • Contribute to AI-assisted operational workflows and internal MCP tooling.
  • Support hosting infrastructure serving thousands of servers and customer workloads across multiple regions.

Requirements

  • 5+ years of experience in SRE, infrastructure, platform, or systems engineering roles operating production systems at scale.
  • Strong Linux fundamentals, including networking, storage, processes, and performance troubleshooting.
  • Proficiency in Python for maintainable, tested automation and tooling.
  • Hands-on experience operating distributed systems and diagnosing issues across service boundaries.
  • Experience with infrastructure-as-code or configuration management using Puppet, Ansible, or an equivalent tool.
  • Experience with production on-call, incident response, reliability ownership, and toil reduction.
  • Clear written and verbal communication skills for documentation, post-incident reviews, and distributed-team collaboration.
  • Experience operating OpenStack, including Nova, Neutron, and Ceph, or comparable cloud infrastructure platforms is preferred.
  • Experience with Docker, Kolla, containerized infrastructure, Ceph or software-defined storage at scale, and cloud migration programs is preferred.
  • Interest in applying AI/LLM tooling to operational workflows is preferred.

Benefits

  • Remote work from home with occasional GoDaddy office visits for team events or meetings.
  • Paid time off and retirement savings benefits, such as 401(k) or pension schemes.
  • Bonus or incentive eligibility and equity grants.
  • Employee stock purchase plan participation.
  • Competitive health benefits.
  • Family-friendly benefits, including parental leave.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

PENN Entertainment 10K-50K Hotels, Restaurants & Leisure

PENN Entertainment is seeking a Senior Site Reliability Engineer to build and operate the cloud infrastructure supporting large-scale sports betting and media platforms across regulated jurisdictions.

Argo CD AWS Bash CI/CD Datadog GCP GitHub Actions GitOps Go Helm Kubernetes Linux PostgreSQL Python Shell Scripting Terraform
16 hours, 53 minutes ago

Staff Site Reliability Engineer, Ads

Reddit 1K-5K Internet Software & Services

Reddit is hiring a Staff Site Reliability Engineer to provide technical leadership for reliability, scalability, and operational excellence across its advertising infrastructure and revenue-critical systems.

Apache Spark ClickHouse GCP Go Kafka Kubernetes
17 hours, 23 minutes ago

Sr Lead Network Reliability Engineer

Coupa Software 1K-5K Internet Software & Services

Coupa is hiring a Sr. Lead Network Development Engineer to scale and operate its global SaaS platform’s cloud networking infrastructure through automation, reliability engineering, and technical leadership.

Ansible AWS Azure Chef DNS Fortinet Go Java Kubernetes Linux Python Ruby TCP/IP Terraform TLS
1 day, 17 hours ago

Senior Monitoring/Observability Architect

Makpar 51-250 Internet Software & Services

Makpar is seeking a Senior Monitoring/Observability Architect to lead enterprise monitoring strategy, architecture, and implementation guidance for a large federal government program.

Datadog Splunk
1 day, 17 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers