Description

  • Operate and maintain all Develocity instances and supporting services.
  • Participate in a follow-the-sun on-call rotation and own incident response across the stack.
  • Drive automation for deployment, upgrades, monitoring, self-healing, and recovery.
  • Build and maintain observability across managed services, including logging, metrics, tracing, and alerting.
  • Collaborate with engineering teams to design reliability into features from the start.
  • Run incident retrospectives and drive follow-up improvements.
  • Own disaster recovery, backups, and business continuity planning.
  • Communicate with customers during incidents and maintenance windows.
  • Optimize performance, resource usage, and infrastructure costs.
  • Help evolve SaaS operations as the company grows.

Requirements

  • 5+ years of experience in SRE, DevOps, or a similar role operating production services at scale.
  • Strong production Kubernetes experience.
  • Cloud infrastructure experience, preferably with AWS (EKS, RDS, S3, EC2).
  • Experience with observability tools such as Prometheus and Grafana, plus Infrastructure as Code with Terraform.
  • Proven incident management and response experience.
  • Knowledge of SRE best practices, including SLAs and SLOs.
  • Scripting proficiency in Python and Bash for automation.
  • Experience with 24/7 on-call rotations.
  • Strong written and verbal English communication skills.
  • Experience operating SaaS platforms at scale (preferred).
  • Familiarity with Develocity (preferred).
  • JVM language experience such as Java or Kotlin (preferred).
  • Disaster recovery planning and execution experience (preferred).
  • Customer-facing incident communication experience (preferred).
  • Experience establishing SRE practices in new or growing teams (preferred).

Benefits

  • US salary range of $150k-$190k.
  • Competitive salaries and equity grants.
  • Remote-first work from anywhere in the PST timezone.
  • A ground-floor role on a new SRE team with real ownership.
  • In-person company offsites and team meetings.
  • A culture that values automation over heroics.
  • Direct customer interaction during incidents and successful recoveries.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

PENN Entertainment 10K-50K Hotels, Restaurants & Leisure

PENN Entertainment is seeking a Senior Site Reliability Engineer to build and operate the cloud infrastructure supporting large-scale sports betting and media platforms across regulated jurisdictions.

Argo CD AWS Bash CI/CD Datadog GCP GitHub Actions GitOps Go Helm Kubernetes Linux PostgreSQL Python Shell Scripting Terraform
17 hours, 14 minutes ago

Staff Site Reliability Engineer, Ads

Reddit 1K-5K Internet Software & Services

Reddit is hiring a Staff Site Reliability Engineer to provide technical leadership for reliability, scalability, and operational excellence across its advertising infrastructure and revenue-critical systems.

Apache Spark ClickHouse GCP Go Kafka Kubernetes
17 hours, 44 minutes ago

Sr Lead Network Reliability Engineer

Coupa Software 1K-5K Internet Software & Services

Coupa is hiring a Sr. Lead Network Development Engineer to scale and operate its global SaaS platform’s cloud networking infrastructure through automation, reliability engineering, and technical leadership.

Ansible AWS Azure Chef DNS Fortinet Go Java Kubernetes Linux Python Ruby TCP/IP Terraform TLS
1 day, 17 hours ago

Senior Monitoring/Observability Architect

Makpar 51-250 Internet Software & Services

Makpar is seeking a Senior Monitoring/Observability Architect to lead enterprise monitoring strategy, architecture, and implementation guidance for a large federal government program.

Datadog Splunk
1 day, 17 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers