Senior Site Reliability Engineer

3 weeks, 4 days ago
Full-time
Senior
DevOps and Infrastructure
Lodgify

Lodgify

Lodgify is a Barcelona-based technology company founded in 2012, offering innovative vacation rental software and website templates. Their all-in-one solution allows property owners to easily create beautiful hospitality websites, accept online booking...

Internet Software & Services
251-1K
Founded 2012
$37M raised

Description

  • Define SLIs, SLOs, error budgets, and reliability targets for platform services.
  • Collaborate with engineering teams to establish observability, production-readiness, and reliability practices.
  • Improve cloud, Kubernetes, and shared infrastructure reliability, scalability, performance, and resilience.
  • Build actionable monitoring with metrics, logs, traces, and golden signals using tools such as Datadog, Prometheus, and Grafana.
  • Automate operational tasks and implement security and operational best practices through code, policies, and guidelines.
  • Develop self-service Internal Developer Platform capabilities using APIs and Kubernetes operators.
  • Improve deployment safety, rollback procedures, and release observability.
  • Strengthen the reliability of databases, caches, queues, and streaming platforms.
  • Participate in on-call rotations, troubleshoot incidents, coordinate response, and lead blameless post-incident reviews.
  • Conduct disaster recovery exercises and identify infrastructure cost and resource-efficiency improvements.

Requirements

  • 7+ years of production experience operating Kubernetes-based platforms and cloud infrastructure.
  • Strong knowledge of SRE practices, including SLIs, SLOs, error budgets, incident response, capacity planning, high availability, backups, and disaster recovery.
  • Experience designing observability and alerting for critical systems using metrics, logs, traces, and golden signals.
  • Ability to troubleshoot complex distributed systems and identify systemic reliability improvements.
  • Ability to write maintainable software, preferably with Python or similar languages, to automate operational work.
  • Experience operating stateful production systems such as relational databases, caches, queues, or streaming platforms.
  • Ability to balance reliability, performance, cost, and delivery speed pragmatically.
  • Experience introducing SRE practices in evolving environments while providing hands-on infrastructure support.
  • Strong collaboration and communication skills across Engineering, Platform, Security, and Product teams.
  • Ability to document clearly, coach teams on production ownership, and drive improvements through completion.

Benefits

  • Flexible remote work options.
  • 25 working days of paid vacation and reduced summer working hours in August.
  • Alan health, dental, and mental health insurance, including coverage for pre-existing conditions.
  • €150 monthly meal allowance and flexible remuneration for meal costs and public transportation.
  • Home office equipment, including a desk, ergonomic chair, and monitor.
  • Free Spanish classes and employee referral rewards.
  • Daily office breakfast, monthly team events, and an international, inclusive workplace.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Service Reliability Engineer

Thoughtworks 10K-50K Professional Services

Thoughtworks is seeking a Senior Site Reliability Engineer to help clients improve infrastructure reliability, resilience, observability, and operational performance through automation and continuous improvement.

AWS Azure CI/CD Datadog ELK Stack GitOps Grafana Kubernetes Microservices New Relic Nomad Python REST API
15 hours, 9 minutes ago

Cloud Site Reliability Engineer

Cadwell 51-250 Health Care Providers & Services

Cadwell is seeking a Cloud Site Reliability Engineer to operate and improve AWS infrastructure supporting healthcare customers and ensure reliable, secure, and compliant hosted neurodiagnostic software environments.

AWS Bash CI/CD Encryption HIPAA JavaScript JSON Python Terraform TypeScript YAML
1 day, 14 hours ago

Site Reliability Engineer

GiveCampus 51-250 Internet Software & Services

GiveCampus is seeking a hands-on Site Reliability Engineer to strengthen the reliability, performance, observability, and operational maturity of its AWS-based fundraising platform in a remote-first U.S. role.

AWS CI/CD CircleCI Datadog GitHub Actions Kubernetes Linux New Relic OpenSearch PostgreSQL Redis Ruby Ruby on Rails Terraform
1 day, 15 hours ago

DevOps / SRE / DevSecOps Engineer (AWS) - Latin America, Remote

Bluelight Consulting 11-50 Internet Software & Services

Bluelight is hiring a DevOps/SRE/DevSecOps professional to operate secure, reliable AWS-based systems and delivery infrastructure for client software projects in a remote consultancy environment.

AWS CI/CD DevSecOps Docker K6 OpenTelemetry PostgreSQL Secrets Management Terraform
2 days, 14 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers