Description

  • Operate and maintain all Develocity production instances and supporting services.
  • Define and evolve SRE standards, practices, and operating models for on-call, incident response, postmortems, and SLOs.
  • Participate in a follow-the-sun on-call rotation and act as an escalation point for severe incidents.
  • Lead incident response and blameless retrospectives, turning learnings into measurable reliability improvements.
  • Set reliability priorities based on risk, customer impact, business goals, SLOs, and error budgets.
  • Identify systemic reliability risks and improve SaaS operations as the platform and customer base grow.
  • Lead architectural and design reviews to improve reliability, scalability, and operability.
  • Drive automation for deployment, upgrades, monitoring, self-healing, recovery, and operational workflows.
  • Build and maintain observability across managed services, including logging, metrics, tracing, and alerting.
  • Own disaster recovery, backups, and business continuity planning and execution.
  • Partner with engineering leadership to balance feature delivery with reliability and operational excellence.
  • Mentor and coach SREs, help onboard new hires, and contribute to hiring.

Requirements

  • 7+ years of experience in SRE, DevOps, or an equivalent role operating production services at scale.
  • Experience leading reliability initiatives across multiple teams or services.
  • Ability to influence technical direction without direct authority.
  • Experience designing and operating systems with SLOs and error budgets.
  • Strong Kubernetes experience in production environments.
  • Cloud infrastructure experience, preferably on AWS (EKS, RDS, S3, EC2).
  • Proficiency with Prometheus, Grafana, and Infrastructure as Code tools such as Terraform.
  • Experience with incident management and response in a 24/7 on-call environment.
  • Scripting proficiency in Python and/or Bash for automation.
  • Strong written and verbal English communication skills.
  • Experience as a founding or early SRE in a growing SaaS organization (preferred).
  • Familiarity with Develocity (preferred).
  • JVM language experience such as Java or Kotlin (preferred).
  • Experience with customer-facing and executive-level incident communications (preferred).

Benefits

  • Competitive salaries and equity grants.
  • Remote-first work from home in Europe (GMT).
  • In-person company offsites and team meetings.
  • A ground-floor role in a new SRE team with real influence over processes and standards.
  • Ownership of production systems used by well-known customers.
  • Direct interaction with customers during incidents and successes.
  • A culture that values automation over heroics.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

PENN Entertainment 10K-50K Hotels, Restaurants & Leisure

PENN Entertainment is seeking a Senior Site Reliability Engineer to build and operate the cloud infrastructure supporting large-scale sports betting and media platforms across regulated jurisdictions.

Argo CD AWS Bash CI/CD Datadog GCP GitHub Actions GitOps Go Helm Kubernetes Linux PostgreSQL Python Shell Scripting Terraform
17 hours, 14 minutes ago

Staff Site Reliability Engineer, Ads

Reddit 1K-5K Internet Software & Services

Reddit is hiring a Staff Site Reliability Engineer to provide technical leadership for reliability, scalability, and operational excellence across its advertising infrastructure and revenue-critical systems.

Apache Spark ClickHouse GCP Go Kafka Kubernetes
17 hours, 44 minutes ago

Sr Lead Network Reliability Engineer

Coupa Software 1K-5K Internet Software & Services

Coupa is hiring a Sr. Lead Network Development Engineer to scale and operate its global SaaS platform’s cloud networking infrastructure through automation, reliability engineering, and technical leadership.

Ansible AWS Azure Chef DNS Fortinet Go Java Kubernetes Linux Python Ruby TCP/IP Terraform TLS
1 day, 17 hours ago

Senior Monitoring/Observability Architect

Makpar 51-250 Internet Software & Services

Makpar is seeking a Senior Monitoring/Observability Architect to lead enterprise monitoring strategy, architecture, and implementation guidance for a large federal government program.

Datadog Splunk
1 day, 17 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers