Site Reliability Engineer

3 weeks, 5 days ago
Full-time
Senior
DevOps and Infrastructure

GoDaddy

GoDaddy is a technology company in the Technology, Information and Internet industry. It describes itself as the world’s largest web services platform and a provider of tools for domains, websites, marketing, email, security, and online business services.

Technology, Information and Internet
5001-10000
Founded 1997

Description

  • Operate, scale, troubleshoot, and improve OpenStack-based compute, networking, and storage services in large-scale production environments.
  • Drive OpenStack migration efforts by designing migration tooling, validation processes, and rollback strategies.
  • Build and maintain Python and Puppet automation to eliminate manual operational work.
  • Improve monitoring, alerting, dashboards, and overall system observability.
  • Participate in shared on-call rotations and lead incident response and blameless post-incident reviews.
  • Review code and technical designs, document systems and runbooks, and mentor SRE I/II engineers.
  • Contribute to AI-assisted operational workflows and internal MCP tooling.
  • Support hosting infrastructure serving thousands of servers and customer workloads across multiple regions.

Requirements

  • 5+ years of experience in SRE, infrastructure, platform, or systems engineering roles operating production systems at scale.
  • Strong Linux fundamentals, including networking, storage, processes, and performance troubleshooting.
  • Proficiency in Python for maintainable, tested automation and tooling.
  • Hands-on experience operating distributed systems and diagnosing issues across service boundaries.
  • Experience with infrastructure-as-code or configuration management using Puppet, Ansible, or an equivalent tool.
  • Experience with production on-call, incident response, reliability ownership, and toil reduction.
  • Clear written and verbal communication skills for documentation, post-incident reviews, and distributed-team collaboration.
  • Experience operating OpenStack, including Nova, Neutron, and Ceph, or comparable cloud infrastructure platforms is preferred.
  • Experience with Docker, Kolla, containerized infrastructure, Ceph or software-defined storage at scale, and cloud migration programs is preferred.
  • Interest in applying AI/LLM tooling to operational workflows is preferred.

Benefits

  • Remote work from home with occasional GoDaddy office visits for team events or meetings.
  • Paid time off and retirement savings benefits, such as 401(k) or pension schemes.
  • Bonus or incentive eligibility and equity grants.
  • Employee stock purchase plan participation.
  • Competitive health benefits.
  • Family-friendly benefits, including parental leave.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Cloud Site Reliability Engineer

Cadwell 51-250 Health Care Providers & Services

Cadwell is seeking a Cloud Site Reliability Engineer to operate and improve AWS infrastructure supporting healthcare customers and ensure reliable, secure, and compliant hosted neurodiagnostic software environments.

AWS Bash CI/CD Encryption HIPAA JavaScript JSON Python Terraform TypeScript YAML
21 hours, 37 minutes ago

Site Reliability Engineer

GiveCampus 51-250 Internet Software & Services

GiveCampus is seeking a hands-on Site Reliability Engineer to strengthen the reliability, performance, observability, and operational maturity of its AWS-based fundraising platform in a remote-first U.S. role.

AWS CI/CD CircleCI Datadog GitHub Actions Kubernetes Linux New Relic OpenSearch PostgreSQL Redis Ruby Ruby on Rails Terraform
22 hours, 7 minutes ago

DevOps / SRE / DevSecOps Engineer (AWS) - Latin America, Remote

Bluelight Consulting 11-50 Internet Software & Services

Bluelight is hiring a DevOps/SRE/DevSecOps professional to operate secure, reliable AWS-based systems and delivery infrastructure for client software projects in a remote consultancy environment.

AWS CI/CD DevSecOps Docker K6 OpenTelemetry PostgreSQL Secrets Management Terraform
1 day, 21 hours ago

DevOps / SRE / DevSecOps Engineer (AWS) - Latin America, Remote

Bluelight Consulting 11-50 Internet Software & Services

Bluelight is hiring a DevOps/SRE/DevSecOps professional to support complex client systems by building secure, reliable, and observable AWS infrastructure and delivery operations.

AWS CI/CD DevSecOps Docker K6 OpenTelemetry PostgreSQL Secrets Management Terraform
1 day, 21 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers