Senior Site Reliability Engineer

9 hours, 28 minutes ago
Full-time
Senior
DevOps and Infrastructure
Lodgify

Lodgify

Lodgify is a Barcelona-based technology company founded in 2012, offering innovative vacation rental software and website templates. Their all-in-one solution allows property owners to easily create beautiful hospitality websites, accept online booking...

Internet Software & Services
251-1K
Founded 2012
$37M raised

Description

  • Define SLIs, SLOs, error budgets, and reliability targets for platform services.
  • Collaborate with engineering teams to establish observability, production-readiness, and reliability practices.
  • Improve cloud, Kubernetes, and shared infrastructure reliability, scalability, performance, and resilience.
  • Build actionable monitoring with metrics, logs, traces, and golden signals using tools such as Datadog, Prometheus, and Grafana.
  • Automate operational tasks and implement security and operational best practices through code, policies, and guidelines.
  • Develop self-service Internal Developer Platform capabilities using APIs and Kubernetes operators.
  • Improve deployment safety, rollback procedures, and release observability.
  • Strengthen the reliability of databases, caches, queues, and streaming platforms.
  • Participate in on-call rotations, troubleshoot incidents, coordinate response, and lead blameless post-incident reviews.
  • Conduct disaster recovery exercises and identify infrastructure cost and resource-efficiency improvements.

Requirements

  • 7+ years of production experience operating Kubernetes-based platforms and cloud infrastructure.
  • Strong knowledge of SRE practices, including SLIs, SLOs, error budgets, incident response, capacity planning, high availability, backups, and disaster recovery.
  • Experience designing observability and alerting for critical systems using metrics, logs, traces, and golden signals.
  • Ability to troubleshoot complex distributed systems and identify systemic reliability improvements.
  • Ability to write maintainable software, preferably with Python or similar languages, to automate operational work.
  • Experience operating stateful production systems such as relational databases, caches, queues, or streaming platforms.
  • Ability to balance reliability, performance, cost, and delivery speed pragmatically.
  • Experience introducing SRE practices in evolving environments while providing hands-on infrastructure support.
  • Strong collaboration and communication skills across Engineering, Platform, Security, and Product teams.
  • Ability to document clearly, coach teams on production ownership, and drive improvements through completion.

Benefits

  • Flexible remote work options.
  • 25 working days of paid vacation and reduced summer working hours in August.
  • Alan health, dental, and mental health insurance, including coverage for pre-existing conditions.
  • €150 monthly meal allowance and flexible remuneration for meal costs and public transportation.
  • Home office equipment, including a desk, ergonomic chair, and monitor.
  • Free Spanish classes and employee referral rewards.
  • Daily office breakfast, monthly team events, and an international, inclusive workplace.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

PointClickCare 1K-5K Health Care Providers & Services

PointClickCare is seeking a Senior Site Reliability Engineer to provide technical leadership and improve the reliability, automation, observability, and operational efficiency of cloud-based healthcare applications.

Agile Ansible AWS Azure C C++ Chef Docker Go Java Kubernetes Linux Perl Puppet Python Ruby TCP/IP Terraform Windows Server
9 hours, 43 minutes ago

AWS - Incident Handler

Caseware 251-1K Internet Software & Services

Caseware is hiring a fully remote Incident Commander in Colombia to lead incident response for its 24/7 SaaS operations, coordinating resolution, communication, root-cause analysis, and post-incident improvements.

AWS JIRA New Relic PagerDuty
2 days, 8 hours ago

Principal Site Reliability Engineer, Platform

Blue River Technology 251-1K Industrial Conglomerates

Blue River Technology, a John Deere company developing AI and robotics for agriculture and construction, is seeking a Principal Site Reliability Engineer to shape and scale the Kubernetes-based platform that enables reliable delivery of autonomous systems and products.

Argo CD AWS CI/CD GitHub Actions Go JavaScript Kubernetes Python Rust Terraform
2 days, 9 hours ago

Senior Site Reliability Engineer - Cloud Platform

GoDaddy 5001-10000 Technology, Information and Internet

GoDaddy’s Global Compute team is seeking a remote infrastructure engineer to operate and scale AWS production infrastructure that supports the company’s engineering teams.

AWS AWS CDK CI/CD CloudFormation GitOps Python
2 days, 9 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers