Site Reliability Engineer (6266)

16 minutes ago
Contract
Junior
DevOps and Infrastructure
Dan.com - a GoDaddy brand

Dan.com - a GoDaddy brand

Dan.com is a GoDaddy brand that offers a wide range of domain-related products and services. With a focus on simplifying the domain buying and selling process, Dan.com provides a user-friendly platform for individuals and businesses to search, register...

Internet Software & Services

Description

  • Develop and maintain infrastructure automation solutions using Ansible for cloud environments.
  • Design, implement, and enhance CI/CD pipelines, testing frameworks, and operational tooling.
  • Troubleshoot complex Linux-based infrastructure and distributed systems issues.
  • Build automation for rapid, repeatable deployment of regional, sovereign, and purpose-built cloud environments.
  • Collaborate with engineering teams, product management, and cross-functional stakeholders on operational improvements.
  • Monitor infrastructure performance and implement enhancements to reduce overhead and improve scalability.
  • Contribute to automation and engineering best practices for reliable cloud platform operations.
  • Participate in internal practice meetings, thought leadership, case studies, and networking events.
  • Work with leadership on career fast-track opportunities and internal practice development.

Requirements

  • 2+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or a related role.
  • Experience developing and maintaining infrastructure automation using Ansible.
  • Experience programming in Ruby and writing automated tests using RSpec or similar frameworks.
  • Experience administering and troubleshooting Linux-based systems and distributed infrastructure environments.
  • Experience designing, implementing, and maintaining CI/CD pipelines, including GitLab CI.
  • Experience supporting large-scale infrastructure environments with hundreds or thousands of systems.
  • Must be eligible to work on FedRAMP projects.
  • Must be a U.S. citizen working from U.S. soil.
  • Bachelor's degree in a relevant field or equivalent work experience.
  • Experience with AWS or other public cloud platforms and hybrid infrastructure environments (preferred).
  • Knowledge of monitoring, observability, and site reliability engineering practices and tools (preferred).
  • Familiarity with Kubernetes concepts and containerized application platforms (preferred).
  • Experience using AI-assisted development tools to improve development and operational productivity (preferred).

Benefits

  • 100% remote work within the United States.
  • 24-month duration.
  • Comprehensive medical benefits.
  • Dental, vision, and life insurance.
  • 401(k) plan with matching.
  • Paid holidays.
  • Networking, career learning, and development programs.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Site Reliability Developer (python/java) / SRE

WatchGuard Technologies 1K-5K Internet Software & Services

WatchGuard is hiring a remote Site Reliability Developer in Spain to support the reliability, security, and operational excellence of its production cloud environments alongside application teams.

Apache Spark AWS Azure CloudFormation Docker Elasticsearch Flink GitHub Go Java Jenkins JIRA Kubernetes Microservices New Relic Python Serverless Terraform
16 minutes ago

Sr. Site Reliability Engineer, tvScientific

Pinterest 5K-10K Internet Software & Services

Pinterest is hiring a Senior Site Reliability Engineer to operate and improve tvScientific’s cloud-native CTV advertising platform on AWS, Kubernetes, and GitOps workflows.

Argo CD AWS Bash CI/CD GCP GitHub Actions GitOps Helm Kubernetes Linux Python Secrets Management Terraform
23 hours, 1 minute ago

Site Reliability Engineering (SRE) Leader

PatSnap 251-1K Internet Software & Services

PatSnap is hiring a Site Reliability Engineering (SRE) Leader to lead its UK SRE team and drive the reliability, scalability, security, and performance of its global SaaS platform.

AWS Docker Kubernetes
1 day ago

SITE RELIABILITY ENGINEER III

Harford County Public Library 51-250 Diversified Consumer Services

Site Reliability Engineer na Stone, atuando no time de Foundation Platform para fortalecer a plataforma interna de tecnologia com foco em observabilidade, automação e estabilidade dos sistemas.

Ansible Argo CD AWS Azure Datadog Docker GCP GitHub Actions Go Grafana Kubernetes Linux Node.js OpenTelemetry Prometheus Python Splunk Terraform
1 day, 23 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers