Site Reliability Engineer, Intermediate to Senior Staff — Infrastructure Platforms

4 weeks, 1 day ago
Full-time
Senior
DevOps and Infrastructure
GitLab

GitLab

GitLab: The comprehensive DevOps platform revolutionizing software development with automation, AI workflows, and essential tools for efficient collaboration.

Internet Software & Services
1K-5K
Founded 2014

Description

  • Keep user-facing services and production systems reliable, scalable, and efficient.
  • Build automation and tooling that reduces toil and replaces manual work with repeatable infrastructure-as-code workflows.
  • Operate and troubleshoot production systems on Kubernetes, including deployments, rollouts, and scaling.
  • Write and maintain infrastructure as code and ship changes safely through CI/CD and GitOps.
  • Participate in on-call rotations, triage alerts, follow and improve runbooks, and escalate appropriately.
  • Contribute to the observability stack by using metrics, logs, and SLOs to detect issues early.
  • Participate in incident response and post-incident reviews, turning lessons learned into improvements in automation and process.
  • Document runbooks, architecture decisions, and reviews so findings become repeatable practices.
  • Lead or contribute to reliability improvements appropriate to the level assigned, from scoped ownership to cross-team strategy.

Requirements

  • Experience keeping production systems reliable using both operations and software engineering practices.
  • Experience building new infrastructure tooling and automation, such as Terraform modules, Kubernetes operators or controllers, or production services written from scratch.
  • Ability to read, debug, and reason about code; most teams use Go, and some use Ruby.
  • Experience with infrastructure as code and with Kubernetes and its ecosystem at a depth appropriate to your level.
  • Hands-on experience with at least one major cloud provider, preferably GCP or AWS.
  • Familiarity with observability practices, including metrics, logging, alerting, and SLOs or SLIs.
  • Comfort participating in on-call and incident response with a structured troubleshooting approach under pressure.
  • Strong written communication skills and ability to operate as a manager-of-one in an async, distributed environment.
  • Track record of using automation, and increasingly AI, to reduce toil and improve team workflows.
  • Alignment with GitLab’s values and willingness to work in accordance with them.

Benefits

  • United States base salary range of $126,400 to $314,400 USD, depending on level and location.
  • Flexible Paid Time Off.
  • Equity compensation and an Employee Stock Purchase Plan.
  • Benefits to support health, finances, and well-being.
  • Growth and Development Fund.
  • Parental leave.
  • All-remote work environment with globally distributed, asynchronous teams.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

PandaDoc 251-1K Internet Software & Services

PandaDoc is hiring a Site Reliability Engineer to help keep its document workflow platform reliable, resilient, and available for customers.

Agile AWS Django Grafana Java Kafka Kubernetes NATS PostgreSQL Python RabbitMQ Spring Boot
16 hours, 16 minutes ago

Senior Site Reliability Engineer

PandaDoc 251-1K Internet Software & Services

PandaDoc is hiring a Site Reliability Engineer to help keep its document workflow platform reliable through incident management, observability, production support, and resilience work across services.

Agile AWS Django Grafana Java Kafka Kubernetes NATS PostgreSQL Python RabbitMQ Spring Boot
16 hours, 16 minutes ago

Senior Site Reliability Engineer

PandaDoc 251-1K Internet Software & Services

PandaDoc is hiring a Site Reliability Engineer to keep its document workflow platform highly available and resilient while supporting production operations and reliability improvements.

Agile AWS Django Grafana Java Kafka Kubernetes NATS PostgreSQL Python RabbitMQ Spring Boot
1 day, 15 hours ago

Senior Site Reliability Engineer

PandaDoc 251-1K Internet Software & Services

PandaDoc is hiring a Site Reliability Engineer to help keep its production document workflow platform reliable, resilient, and low-downtime for customers.

AWS Django Grafana Java Kafka Kubernetes NATS PostgreSQL Python RabbitMQ Spring Boot
1 day, 16 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers