Description

  • Operate and maintain all Develocity production instances and supporting services.
  • Define and evolve SRE standards, practices, and operating models for on-call, incident response, postmortems, and SLOs.
  • Participate in a follow-the-sun on-call rotation and act as an escalation point for severe incidents.
  • Lead incident response and blameless retrospectives, turning learnings into measurable reliability improvements.
  • Set reliability priorities based on risk, customer impact, business goals, SLOs, and error budgets.
  • Identify systemic reliability risks and improve SaaS operations as the platform and customer base grow.
  • Lead architectural and design reviews to improve reliability, scalability, and operability.
  • Drive automation for deployment, upgrades, monitoring, self-healing, recovery, and operational workflows.
  • Build and maintain observability across managed services, including logging, metrics, tracing, and alerting.
  • Own disaster recovery, backups, and business continuity planning and execution.
  • Partner with engineering leadership to balance feature delivery with reliability and operational excellence.
  • Mentor and coach SREs, help onboard new hires, and contribute to hiring.

Requirements

  • 7+ years of experience in SRE, DevOps, or an equivalent role operating production services at scale.
  • Experience leading reliability initiatives across multiple teams or services.
  • Ability to influence technical direction without direct authority.
  • Experience designing and operating systems with SLOs and error budgets.
  • Strong Kubernetes experience in production environments.
  • Cloud infrastructure experience, preferably on AWS (EKS, RDS, S3, EC2).
  • Proficiency with Prometheus, Grafana, and Infrastructure as Code tools such as Terraform.
  • Experience with incident management and response in a 24/7 on-call environment.
  • Scripting proficiency in Python and/or Bash for automation.
  • Strong written and verbal English communication skills.
  • Experience as a founding or early SRE in a growing SaaS organization (preferred).
  • Familiarity with Develocity (preferred).
  • JVM language experience such as Java or Kotlin (preferred).
  • Experience with customer-facing and executive-level incident communications (preferred).

Benefits

  • Competitive salaries and equity grants.
  • Remote-first work from home in Europe (GMT).
  • In-person company offsites and team meetings.
  • A ground-floor role in a new SRE team with real influence over processes and standards.
  • Ownership of production systems used by well-known customers.
  • Direct interaction with customers during incidents and successes.
  • A culture that values automation over heroics.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Site Reliability Engineer

GiveCampus 51-250 Internet Software & Services

GiveCampus is seeking a hands-on Site Reliability Engineer to strengthen the reliability, performance, observability, and operational maturity of its AWS-based fundraising platform in a remote-first U.S. role.

AWS CI/CD CircleCI Datadog GitHub Actions Kubernetes Linux New Relic OpenSearch PostgreSQL Redis Ruby Ruby on Rails Terraform
14 minutes ago

DevOps / SRE / DevSecOps Engineer (AWS) - Latin America, Remote

Bluelight Consulting 11-50 Internet Software & Services

Bluelight is hiring a DevOps/SRE/DevSecOps professional to operate secure, reliable AWS-based systems and delivery infrastructure for client software projects in a remote consultancy environment.

AWS CI/CD DevSecOps Docker K6 OpenTelemetry PostgreSQL Secrets Management Terraform
23 hours, 44 minutes ago

DevOps / SRE / DevSecOps Engineer (AWS) - Latin America, Remote

Bluelight Consulting 11-50 Internet Software & Services

Bluelight is hiring a DevOps/SRE/DevSecOps professional to support complex client systems by building secure, reliable, and observable AWS infrastructure and delivery operations.

AWS CI/CD DevSecOps Docker K6 OpenTelemetry PostgreSQL Secrets Management Terraform
23 hours, 44 minutes ago

DevOps / SRE / DevSecOps Engineer (AWS) - Latin America, Remote

Bluelight Consulting 11-50 Internet Software & Services

Bluelight is hiring a DevOps/SRE professional to support complex, security-sensitive client systems by building reliable AWS infrastructure, deployment processes, and operational practices within its remote software consultancy.

AWS CI/CD DevSecOps Docker K6 OpenTelemetry PostgreSQL Secrets Management Terraform
23 hours, 44 minutes ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers