Staff Site Reliability Engineer

3 hours, 17 minutes ago
Full-time
Lead
DevOps and Infrastructure

Description

  • Operate and maintain all Develocity production instances and supporting services.
  • Define and evolve SRE standards, practices, and operating models for on-call, incident response, postmortems, and SLOs.
  • Participate in a follow-the-sun on-call rotation and act as an escalation point for severe incidents.
  • Lead incident response and blameless retrospectives, turning learnings into measurable reliability improvements.
  • Set reliability priorities based on risk, customer impact, business goals, SLOs, and error budgets.
  • Identify systemic reliability risks and improve SaaS operations as the platform and customer base grow.
  • Lead architectural and design reviews to improve reliability, scalability, and operability.
  • Drive automation for deployment, upgrades, monitoring, self-healing, recovery, and operational workflows.
  • Build and maintain observability across managed services, including logging, metrics, tracing, and alerting.
  • Own disaster recovery, backups, and business continuity planning and execution.
  • Partner with engineering leadership to balance feature delivery with reliability and operational excellence.
  • Mentor and coach SREs, help onboard new hires, and contribute to hiring.

Requirements

  • 7+ years of experience in SRE, DevOps, or an equivalent role operating production services at scale.
  • Experience leading reliability initiatives across multiple teams or services.
  • Ability to influence technical direction without direct authority.
  • Experience designing and operating systems with SLOs and error budgets.
  • Strong Kubernetes experience in production environments.
  • Cloud infrastructure experience, preferably on AWS (EKS, RDS, S3, EC2).
  • Proficiency with Prometheus, Grafana, and Infrastructure as Code tools such as Terraform.
  • Experience with incident management and response in a 24/7 on-call environment.
  • Scripting proficiency in Python and/or Bash for automation.
  • Strong written and verbal English communication skills.
  • Experience as a founding or early SRE in a growing SaaS organization (preferred).
  • Familiarity with Develocity (preferred).
  • JVM language experience such as Java or Kotlin (preferred).
  • Experience with customer-facing and executive-level incident communications (preferred).

Benefits

  • Competitive salaries and equity grants.
  • Remote-first work from home in Europe (GMT).
  • In-person company offsites and team meetings.
  • A ground-floor role in a new SRE team with real influence over processes and standards.
  • Ownership of production systems used by well-known customers.
  • Direct interaction with customers during incidents and successes.
  • A culture that values automation over heroics.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Platform Engineering Manager

Prolific 51-250 Professional Services

Prolific is hiring a Platform Engineering Manager to lead its Cloud Platform and SRE teams, owning the technical foundation, reliability, and scalability of the infrastructure that supports its AI-focused human data platform.

Argo CD AWS Celery CircleCI Datadog DynamoDB Elasticsearch GCP GitHub Actions GitOps JavaScript Kubernetes MongoDB PostgreSQL Python Serverless Terraform TypeScript
2 hours, 17 minutes ago

Site Reliability Engineer

Axle Informatics 51-250 Pharmaceuticals

Axle is hiring a Site Reliability Engineer to modernize and unify its multi-cloud platform supporting scientific and clinical programs across AWS, Azure, and GCP.

Ansible AWS Azure Bash CentOS Chef CI/CD Docker GCP GitHub Actions Grafana Java Jenkins Kubernetes Linux MLOps .NET Node.js OpenTelemetry PowerShell Prometheus Puppet Python R Splunk SQL Terraform Ubuntu
1 day, 2 hours ago

Site Reliability Developer (python/java) / SRE

WatchGuard Technologies 1K-5K Internet Software & Services

WatchGuard is hiring a remote Site Reliability Developer in Spain to support the reliability, security, and operational excellence of its production cloud environments alongside application teams.

Apache Spark AWS Azure CloudFormation Docker Elasticsearch Flink GitHub Go Java Jenkins JIRA Kubernetes Microservices New Relic Python Serverless Terraform
1 day, 3 hours ago

Site Reliability Engineer (6266)

Dan.com - a GoDaddy brand Internet Software & Services

itD is hiring a Site Reliability Engineer for a 100% remote U.S.-based contract focused on automating and improving the reliability of large-scale cloud infrastructure and production environments.

Ansible AWS CI/CD GitLab CI Kubernetes Linux RSpec Ruby
1 day, 3 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers