Staff Software Engineer - Databases SRE | Spain | Remote

1 month, 4 weeks ago
Grafana

Grafana

Grafana is the open observability platform providing analytics, monitoring, and visualization solutions with a focus on user control and cost efficiency.

IT Services
1K-5K
Founded 2014
$535M raised

Description

  • Partner closely with embedded product engineering squads to support cloud database reliability.
  • Own production reliability for high-SLA and complex customer environments.
  • Design and implement automation to scale reliability practices and reduce toil.
  • Define, review, and evolve per-tenant SLOs and reliability models.
  • Proactively investigate and reduce SLO burn to prevent repeat incidents.
  • Serve as a primary escalation point and participate in on-call and incident response.
  • Lead customer-impacting incident investigations, post-incident reviews, and follow-up actions.
  • Contribute to design documents, pull request reviews, and product design decisions.
  • Improve observability, alert quality, self-healing, and auto-scaling in production systems.
  • Teach Site Reliability Engineering practices and communicate them early in feature development.

Requirements

  • 8+ years of engineering experience, including 4+ years in SRE, CRE, or production engineering.
  • Strong Kubernetes experience in AWS, GCP, or Azure.
  • Familiarity with infrastructure-as-code tools such as Helm, Terraform, or Jsonnet.
  • Strong technical leadership experience, including mentoring engineers and leading projects.
  • Experience operating multi-tenant systems in production.
  • Strong experience designing and implementing SLOs.
  • Experience with one or more programming languages such as Go, Python, or Java.
  • Knowledge of Linux internals plus networking, cloud storage, and scaling concepts.
  • Experience with incident response, including writing high-quality PIRs or post-mortems.
  • Ability to reason about performance, scaling, and failure modes.
  • Comfort working with high autonomy and self-direction in an engineering team.
  • Strong collaboration skills for partnering deeply with product engineering teams.
  • Preferred: formal customer reliability engineering experience.
  • Preferred: intellectually curious, transparent, action-oriented, and kind.

Benefits

  • Base compensation in Spain of €94,025 - €112,830.
  • Equity is included in the compensation package.
  • Bonus eligibility may apply.
  • 100% remote work with a global team.
  • Global annual leave policy of 30 days per year.
  • 3 days of annual leave reserved for Grafana Shutdown Days.
  • In-person onboarding for new hires.
  • Country-specific pay range and benefits for other locations.
  • Equal opportunity employer with a commitment to diversity and inclusion.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Appian Consultant (FMS)

Horizon Industries Limited is hiring a remote, full-time Appian Consultant to support Foreign Military Sales through Appian development, solution delivery, and growing technical ownership.

Agile MySQL Oracle PostgreSQL SQL
16 hours, 17 minutes ago

Principal Software Engineer, Pricing Platform

Upstart 1K-5K Banks

Upstart is seeking a Principal Software Engineer to provide cross-functional technical leadership and evolve the scalable pricing platform that converts machine-learning underwriting outputs into real-time pricing decisions across its lending products.

AWS Celery Datadog Django FastAPI Go gRPC Helm Java Kotlin Kubernetes Machine Learning Microservices PostgreSQL Prototyping Python Rust Spring Boot Statistics
16 hours, 17 minutes ago

Staff Software Engineer, Backend (Identity International)

Affirm 1K-5K Diversified Financial Services

Affirm is hiring a senior backend engineering leader for its International Expansion Unit in the UK to set technical strategy and evolve global identity systems supporting compliant market expansion, customer trust, and business growth.

Apache Spark AWS Kotlin Kubernetes MySQL Python
16 hours, 32 minutes ago

Senior Site Reliability Engineer

PENN Entertainment 10K-50K Hotels, Restaurants & Leisure

PENN Entertainment is seeking a Senior Site Reliability Engineer to build and operate the cloud infrastructure supporting large-scale sports betting and media platforms across regulated jurisdictions.

Argo CD AWS Bash CI/CD Datadog GCP GitHub Actions GitOps Go Helm Kubernetes Linux PostgreSQL Python Shell Scripting Terraform
16 hours, 47 minutes ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers