Senior Site Reliability Engineer

3 months ago
Full-time
Senior
DevOps and Infrastructure
Zipdev

Zipdev

Zipdev specializes in connecting businesses with highly skilled remote talent from Latin America, offering a cost-effective solution for hiring software developers and virtual assistants.

Professional Services
51-250
Founded 2014

Description

  • Support observability implementation using Datadog and/or Azure Monitor/App Insights, including SLOs, alert rules, and synthetic checks.
  • Participate in a PagerDuty on-call rotation and handle escalations, incidents, and documentation.
  • Build and maintain operational runbooks for incident response, rollback, and recovery.
  • Contribute to deployment automation using blue/green or canary patterns and Infrastructure as Code.
  • Work across Azure SQL and Cosmos DB environments to support performance and cost optimization.
  • Collaborate closely with US-based engineers during overlapping working hours.
  • Own production operations as a full-time SRE rather than working in a support or ticket-queue role.

Requirements

  • 5+ years of experience in SRE, DevOps, or cloud infrastructure roles.
  • Strong hands-on experience with Microsoft Azure, including Azure SQL, Cosmos DB, Container Apps, and App Service.
  • Experience with observability tools such as Datadog, Azure Monitor, or similar, plus on-call and incident response experience.
  • Familiarity with Infrastructure as Code, with Terraform preferred.
  • Strong written and spoken English for daily communication with US-based teammates and client stakeholders.
  • Availability with meaningful overlap with US Eastern or Mountain time zones.
  • Experience working in HIPAA-regulated environments, including handling PHI under a Business Associate Agreement and using least-privilege, audited access controls.
  • Willingness to complete a healthcare-industry-standard background check before production access.

Benefits

  • Remote work.
  • 10 business days of vacation per year.
  • 5 national holidays and 5 company holidays per year.
  • Parental leave.
  • Health care reimbursement.
  • Active lifestyle reimbursement.
  • Quarterly home office reimbursement.
  • Payroll deduction purchase plans.
  • Longevity bonus.
  • Continuous learning bonus.
  • Access to training and professional development platforms.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

OMS Application Engineer

Warner Music Group is seeking an OMS Application Engineer to maintain, upgrade, and support the technical systems powering its global commerce ecosystem, with a focus on reliability, scalability, and long-term stability.

AWS CI/CD GitHub Actions Java Microservices
1 day, 10 hours ago

[Job-32080] Senior SRE / Cloud Engineer (Pessoa Engenheira de Plataforma), Brazil

CI&T 5K-10K Internet Software & Services

A CI&T busca uma pessoa Engenheira de Plataforma Sênior para projetar e operar a infraestrutura OCI que sustenta agentes de IA, garantindo alta disponibilidade e baixa latência.

Generative AI Kafka Kubernetes Terraform
1 day, 10 hours ago

Sr. Sustaining and Forward Deployed Engineer

Abacus Insights 51-250 Insurance

Abacus Insights is seeking a Senior Site Reliability Engineer – Forward Deployed to operate and improve its AWS- and Databricks-based healthcare data platform while resolving complex production issues and supporting customers.

Apache Spark AWS CI/CD Databricks Kubernetes Python Snowflake
2 days, 11 hours ago

Site Reliability Engineer

Raydar Professional Services

A Site Reliability Engineer at a remote conversational AI company will build and operate the CI/CD, developer platforms, observability, and agent infrastructure supporting reliable voice-agent products at enterprise scale.

CI/CD Helm Kubernetes Node.js Python Terraform TypeScript
3 days, 11 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers