Senior Site Reliability Specialist II

8 hours, 40 minutes ago
Full-time
Senior
Software Development
Everbridge

Everbridge

Everbridge provides a comprehensive software platform that automates and enhances organizations' responses to critical events, ensuring the safety of individuals and the continuity of business operations during emergencies such as natural disasters, cy...

Internet Software & Services
1K-5K
Founded 2002

Description

  • Build platform capabilities, automation, paved roads, and self-service tools that enable reliable software delivery.
  • Lead cross-functional initiatives across cloud infrastructure, Kubernetes, observability, networking, automation, and developer platforms.
  • Design and implement solutions that improve platform availability, scalability, performance, and resilience.
  • Improve monitoring, alerting, telemetry, incident response, disaster recovery, capacity planning, and production readiness.
  • Use production data and reliability metrics to identify systemic improvements and reduce operational toil.
  • Review architectures and establish engineering standards focused on reliability, recoverability, scalability, and operational excellence.
  • Coach engineering teams on SLOs, error budgets, observability, incident response, and reliable service ownership.
  • Participate in on-call rotations, lead high-severity incident responses, and facilitate blameless post-incident reviews.
  • Drive corrective actions to completion and share knowledge through documentation, design reviews, and collaborative problem solving.

Requirements

  • Experience designing and operating complex production systems.
  • Experience with cloud infrastructure and cloud-native architectures.
  • Experience with distributed systems and container platforms, including Kubernetes.
  • Experience with Infrastructure as Code, automation, CI/CD, and software delivery practices.
  • Experience with observability, monitoring, logging, and telemetry.
  • Knowledge of incident response and operational excellence practices.
  • Understanding of reliability engineering principles, including SLOs, SLIs, capacity planning, and performance optimization.
  • Ability to write software or automation in one or more modern programming languages.
  • Strong Linux and networking fundamentals.
  • Experience in regulated environments such as FedRAMP, DoD IL4/IL5, SOC 2, or ISO 27001 is preferred.

Benefits

  • Estimated salary of $145,000–$177,000, with potential variable compensation.
  • Healthcare, dental, mental health, disability income, and life and AD&D insurance benefits.
  • Parental planning benefits.
  • 401(k) plan with company match.
  • Paid time off.
  • Fitness reimbursements.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Site Reliability Engineer (SRE)

Rocket.net 11-50 IT Services

Rocket.net is seeking a Site Reliability Engineer to maintain the reliability and performance of its hosting platform while resolving complex infrastructure issues and providing advanced support to customers.

Apache Bash CDN Cloudflare Datadog DNS Linux MariaDB MySQL Nginx Redis SSH WAF WordPress
8 hours, 40 minutes ago

Senior Site Reliability Engineer

GoDaddy 5001-10000 Technology, Information and Internet

GoDaddy is hiring a Senior Site Reliability Engineer to support the Commerce ecosystem remotely by improving the reliability, scalability, security, and operation of business-critical production platforms.

Ansible AWS AWS CDK CI/CD CloudFormation Go Kubernetes Linux Pulumi Python SaltStack Terraform TypeScript
1 day, 8 hours ago

Senior Site Reliability Engineer

PENN Entertainment 10K-50K Hotels, Restaurants & Leisure

PENN Entertainment is seeking a Senior Site Reliability Engineer to build and operate the cloud infrastructure supporting large-scale sports betting and media platforms across regulated jurisdictions.

Argo CD AWS Bash CI/CD Datadog GCP GitHub Actions GitOps Go Helm Kubernetes Linux PostgreSQL Python Shell Scripting Terraform
2 days, 7 hours ago

Staff Site Reliability Engineer, Ads

Reddit 1K-5K Internet Software & Services

Reddit is hiring a Staff Site Reliability Engineer to provide technical leadership for reliability, scalability, and operational excellence across its advertising infrastructure and revenue-critical systems.

Apache Spark ClickHouse GCP Go Kafka Kubernetes
2 days, 8 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers