Senior Site Reliability Engineer - Cloud Platform

2 days, 10 hours ago
Full-time
Senior
Software Development

GoDaddy

GoDaddy is a technology company in the Technology, Information and Internet industry. It describes itself as the world’s largest web services platform and a provider of tools for domains, websites, marketing, email, security, and online business services.

Technology, Information and Internet
5001-10000
Founded 1997

Description

  • Operate and scale AWS production infrastructure across GoDaddy AWS organizations.
  • Design, build, and maintain cloud platform capabilities using Python, CloudFormation, AWS CDK, and automation practices.
  • Drive cloud cost optimization initiatives and improve observability through monitoring, alerting, dashboards, and operational tooling.
  • Participate in on-call rotations, lead incident response, and implement reliability improvements through blameless post-incident reviews.
  • Support AWS initiatives involving networking, identity, governance, and multi-account architecture.
  • Review code and designs, maintain documentation and operational runbooks, and mentor fellow engineers.
  • Use AI-assisted tooling to improve productivity, accelerate automation, and reduce operational toil.

Requirements

  • 5+ years of experience in Site Reliability Engineering, Platform Engineering, Infrastructure Engineering, or a similar role supporting production AWS environments.
  • Strong Python development skills for building automation, services, or platform tooling.
  • Hands-on Infrastructure as Code experience with CloudFormation and/or AWS CDK.
  • Experience operating AWS services including IAM, VPC networking, EC2, Lambda, and managed storage or database services.
  • Strong Linux systems administration, troubleshooting, incident response, and production operations experience.
  • Experience operating multi-account AWS environments or AWS Organizations (preferred).
  • Experience with AWS networking technologies such as Transit Gateway, IPAM, or BYOIP (preferred).
  • Cloud cost optimization or FinOps experience, CI/CD and GitOps deployment experience, or exposure to policy-as-code and cloud governance tooling (preferred).
  • Familiarity with AI-assisted engineering workflows and tooling (preferred).

Benefits

  • Remote work from home with occasional office visits.
  • CAD $107,000–$161,000 salary range.
  • Potential eligibility for a discretionary bonus of 10% of base salary.
  • Potential eligibility for GoDaddy’s equity plan and employee stock purchase plan.
  • Health, dental, vision, life, and critical illness insurance.
  • Paid holidays, sick and personal time, wellness days, and parental leave.
  • Retirement savings program and employee assistance program.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

Lodgify 251-1K Internet Software & Services

Lodgify, a Barcelona-based vacation-rental technology company, is hiring a Senior Site Reliability Engineer to improve the reliability, scalability, observability, and operational ownership of its cloud platform and critical product services.

Datadog Grafana Kubernetes Microservices Prometheus Python
10 hours, 9 minutes ago

Senior Site Reliability Engineer

PointClickCare 1K-5K Health Care Providers & Services

PointClickCare is seeking a Senior Site Reliability Engineer to provide technical leadership and improve the reliability, automation, observability, and operational efficiency of cloud-based healthcare applications.

Agile Ansible AWS Azure C C++ Chef Docker Go Java Kubernetes Linux Perl Puppet Python Ruby TCP/IP Terraform Windows Server
10 hours, 24 minutes ago

AWS - Incident Handler

Caseware 251-1K Internet Software & Services

Caseware is hiring a fully remote Incident Commander in Colombia to lead incident response for its 24/7 SaaS operations, coordinating resolution, communication, root-cause analysis, and post-incident improvements.

AWS JIRA New Relic PagerDuty
2 days, 9 hours ago

Principal Site Reliability Engineer, Platform

Blue River Technology 251-1K Industrial Conglomerates

Blue River Technology, a John Deere company developing AI and robotics for agriculture and construction, is seeking a Principal Site Reliability Engineer to shape and scale the Kubernetes-based platform that enables reliable delivery of autonomous systems and products.

Argo CD AWS CI/CD GitHub Actions Go JavaScript Kubernetes Python Rust Terraform
2 days, 10 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers