Staff Site Reliability Engineer

4 weeks, 2 days ago
Full-time
Lead
DevOps and Infrastructure
Filevine

Filevine

Filevine is a top legal tech company revolutionizing legal work with AI-powered case management software, empowering law firms to streamline operations and enhance client services.

Specialized Consumer Services
251-1K
Founded 2015
$226M raised

Description

  • Define and execute the technical strategy for observability, alerting, platform infrastructure, and operational excellence.
  • Lead the evolution of reliable, scalable, secure, and efficient cloud platforms and distributed systems.
  • Champion SLIs, SLOs, error budgets, capacity planning, operational readiness, and automation across the service lifecycle.
  • Lead complex production incidents and drive post-incident learnings into permanent engineering improvements.
  • Build self-service platform capabilities that reduce toil, improve safety, and increase engineering velocity.
  • Mentor engineers and serve as a trusted technical authority for long-term reliability and platform direction.
  • Partner with engineering leadership and the Reliability Architect on major technical decisions and reliability strategy.
  • Shape on-call strategy, tooling, and culture to make production support more sustainable.
  • Drive the use of AI and machine learning in observability, anomaly detection, incident response, automated remediation, and resource optimization.

Requirements

  • 12+ years of experience in software engineering, infrastructure, platform engineering, or SRE.
  • 6+ years of experience in SRE.
  • 3+ years leading complex, cross-functional technical initiatives for distributed production systems.
  • Expert-level depth in observability and platform infrastructure.
  • Broad expertise in incident response, capacity planning, automation, and reliability engineering.
  • Advanced experience with a major container-orchestration platform, preferably Kubernetes.
  • Experience with an observability platform such as New Relic, Datadog, or equivalent.
  • Strong software engineering ability in Python, Go, Bash, or another general-purpose language.
  • Experience building production tooling, automation, or platform capabilities.
  • Experience in a regulated environment such as FedRAMP, CJIS, HIPAA, SOC 2, or PCI is strongly preferred.

Benefits

  • Medical, dental, and vision insurance for full-time employees.
  • Competitive and fair pay.
  • Maternity and paternity leave for full-time employees.
  • Short- and long-term disability coverage.
  • Opportunity to learn from a dedicated leadership team.
  • A dynamic, rapidly growing company environment.
  • Top-of-the-line company swag.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Specialist II

Everbridge 1K-5K Internet Software & Services

Everbridge is hiring a Senior Site Reliability Engineer II to build resilient platforms and improve the reliability, scalability, and operational readiness of systems supporting critical-event communications.

CI/CD Kubernetes Linux
9 hours, 15 minutes ago

Site Reliability Engineer (SRE)

Rocket.net 11-50 IT Services

Rocket.net is seeking a Site Reliability Engineer to maintain the reliability and performance of its hosting platform while resolving complex infrastructure issues and providing advanced support to customers.

Apache Bash CDN Cloudflare Datadog DNS Linux MariaDB MySQL Nginx Redis SSH WAF WordPress
9 hours, 15 minutes ago

Senior Site Reliability Engineer

GoDaddy 5001-10000 Technology, Information and Internet

GoDaddy is hiring a Senior Site Reliability Engineer to support the Commerce ecosystem remotely by improving the reliability, scalability, security, and operation of business-critical production platforms.

Ansible AWS AWS CDK CI/CD CloudFormation Go Kubernetes Linux Pulumi Python SaltStack Terraform TypeScript
1 day, 8 hours ago

Senior Site Reliability Engineer

PENN Entertainment 10K-50K Hotels, Restaurants & Leisure

PENN Entertainment is seeking a Senior Site Reliability Engineer to build and operate the cloud infrastructure supporting large-scale sports betting and media platforms across regulated jurisdictions.

Argo CD AWS Bash CI/CD Datadog GCP GitHub Actions GitOps Go Helm Kubernetes Linux PostgreSQL Python Shell Scripting Terraform
2 days, 8 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers