Staff Software Engineer - Databases SRE | Sweden | Remote

1 month, 4 weeks ago
Grafana

Grafana

Grafana is the open observability platform providing analytics, monitoring, and visualization solutions with a focus on user control and cost efficiency.

IT Services
1K-5K
Founded 2014
$535M raised

Description

  • Partner closely with embedded product engineering squads to support reliability work across Mimir, Loki, Tempo, and Pyroscope.
  • Own production reliability for high-SLA and complex customer environments.
  • Design and implement automation to scale reliability practices and eliminate toil.
  • Define, review, and evolve per-tenant SLOs and reliability models.
  • Proactively reduce SLO burn and investigate reliability issues before they become repeat incidents.
  • Serve as a primary escalation point and participate in on-call and incident response.
  • Lead customer-impacting incident investigations, post-incident reviews, and customer communications when needed.
  • Contribute to design docs, PR reviews, and code reviews.
  • Influence feature design to improve production scalability, operability, and fault tolerance.
  • Improve monitoring quality, self-healing, alerting, and observability for customer environments.
  • Collaborate with engineering leaders on product strategy, roadmaps, and technical designs.
  • Teach SRE best practices and help teams apply them early in feature development.

Requirements

  • 8+ years of engineering experience, including 4+ years in SRE, CRE, or production engineering.
  • Strong Kubernetes experience in AWS, GCP, or Azure.
  • Experience with infrastructure-as-code tools such as Helm, Terraform, or Jsonnet.
  • Experience leading technical projects and mentoring other engineers.
  • Experience operating multi-tenant systems in production.
  • Strong experience designing and implementing SLOs.
  • Experience with one or more programming languages such as Go, Python, or Java.
  • Knowledge of Linux internals, networking, cloud storage, and scaling concepts.
  • Excellent problem-solving and troubleshooting skills.
  • Experience participating in blame-free incident response and writing high-quality PIRs or post-mortems.
  • Ability to reason about performance, scaling, and failure modes.
  • Comfort working with high-autonomy, self-directed engineering teams.
  • Strong ability to partner deeply with product engineering teams.
  • Customer reliability engineering experience is strongly preferred.
  • Intellectual curiosity, transparency, bias for action, and kindness are highly valued.

Benefits

  • Base compensation in Sweden of SEK 878,578 to SEK 1,054,294, depending on level, experience, and skillset.
  • Equity and bonus eligibility, where applicable.
  • 100% remote work with a global team.
  • Global annual leave policy of 30 days per year.
  • 3 days of annual leave reserved for Grafana Shutdown Days.
  • In-person onboarding for new hires.
  • Career growth pathways and development opportunities.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Appian Consultant (FMS)

Horizon Industries Limited is hiring a remote, full-time Appian Consultant to support Foreign Military Sales through Appian development, solution delivery, and growing technical ownership.

Agile MySQL Oracle PostgreSQL SQL
16 hours, 18 minutes ago

Principal Software Engineer, Pricing Platform

Upstart 1K-5K Banks

Upstart is seeking a Principal Software Engineer to provide cross-functional technical leadership and evolve the scalable pricing platform that converts machine-learning underwriting outputs into real-time pricing decisions across its lending products.

AWS Celery Datadog Django FastAPI Go gRPC Helm Java Kotlin Kubernetes Machine Learning Microservices PostgreSQL Prototyping Python Rust Spring Boot Statistics
16 hours, 18 minutes ago

Staff Software Engineer, Backend (Identity International)

Affirm 1K-5K Diversified Financial Services

Affirm is hiring a senior backend engineering leader for its International Expansion Unit in the UK to set technical strategy and evolve global identity systems supporting compliant market expansion, customer trust, and business growth.

Apache Spark AWS Kotlin Kubernetes MySQL Python
16 hours, 33 minutes ago

Senior Site Reliability Engineer

PENN Entertainment 10K-50K Hotels, Restaurants & Leisure

PENN Entertainment is seeking a Senior Site Reliability Engineer to build and operate the cloud infrastructure supporting large-scale sports betting and media platforms across regulated jurisdictions.

Argo CD AWS Bash CI/CD Datadog GCP GitHub Actions GitOps Go Helm Kubernetes Linux PostgreSQL Python Shell Scripting Terraform
16 hours, 48 minutes ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers