Staff Site Reliability Engineer

1 month, 2 weeks ago
Full-time
Lead
DevOps and Infrastructure
Filevine

Filevine

Filevine is a top legal tech company revolutionizing legal work with AI-powered case management software, empowering law firms to streamline operations and enhance client services.

Specialized Consumer Services
251-1K
Founded 2015
$226M raised

Description

  • Define and execute the technical strategy for observability, alerting, platform infrastructure, and operational excellence.
  • Lead the evolution of reliable, scalable, secure, and efficient cloud platforms and distributed systems.
  • Champion SLIs, SLOs, error budgets, capacity planning, operational readiness, and automation across the service lifecycle.
  • Lead complex production incidents and drive post-incident learnings into permanent engineering improvements.
  • Build self-service platform capabilities that reduce toil, improve safety, and increase engineering velocity.
  • Mentor engineers and serve as a trusted technical authority for long-term reliability and platform direction.
  • Partner with engineering leadership and the Reliability Architect on major technical decisions and reliability strategy.
  • Shape on-call strategy, tooling, and culture to make production support more sustainable.
  • Drive the use of AI and machine learning in observability, anomaly detection, incident response, automated remediation, and resource optimization.

Requirements

  • 12+ years of experience in software engineering, infrastructure, platform engineering, or SRE.
  • 6+ years of experience in SRE.
  • 3+ years leading complex, cross-functional technical initiatives for distributed production systems.
  • Expert-level depth in observability and platform infrastructure.
  • Broad expertise in incident response, capacity planning, automation, and reliability engineering.
  • Advanced experience with a major container-orchestration platform, preferably Kubernetes.
  • Experience with an observability platform such as New Relic, Datadog, or equivalent.
  • Strong software engineering ability in Python, Go, Bash, or another general-purpose language.
  • Experience building production tooling, automation, or platform capabilities.
  • Experience in a regulated environment such as FedRAMP, CJIS, HIPAA, SOC 2, or PCI is strongly preferred.

Benefits

  • Medical, dental, and vision insurance for full-time employees.
  • Competitive and fair pay.
  • Maternity and paternity leave for full-time employees.
  • Short- and long-term disability coverage.
  • Opportunity to learn from a dedicated leadership team.
  • A dynamic, rapidly growing company environment.
  • Top-of-the-line company swag.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Service Reliability Engineer

Thoughtworks 10K-50K Professional Services

Thoughtworks is seeking a Senior Site Reliability Engineer to help clients improve infrastructure reliability, resilience, observability, and operational performance through automation and continuous improvement.

AWS Azure CI/CD Datadog ELK Stack GitOps Grafana Kubernetes Microservices New Relic Nomad Python REST API
10 hours, 53 minutes ago

Cloud Site Reliability Engineer

Cadwell 51-250 Health Care Providers & Services

Cadwell is seeking a Cloud Site Reliability Engineer to operate and improve AWS infrastructure supporting healthcare customers and ensure reliable, secure, and compliant hosted neurodiagnostic software environments.

AWS Bash CI/CD Encryption HIPAA JavaScript JSON Python Terraform TypeScript YAML
1 day, 10 hours ago

Site Reliability Engineer

GiveCampus 51-250 Internet Software & Services

GiveCampus is seeking a hands-on Site Reliability Engineer to strengthen the reliability, performance, observability, and operational maturity of its AWS-based fundraising platform in a remote-first U.S. role.

AWS CI/CD CircleCI Datadog GitHub Actions Kubernetes Linux New Relic OpenSearch PostgreSQL Redis Ruby Ruby on Rails Terraform
1 day, 10 hours ago

DevOps / SRE / DevSecOps Engineer (AWS) - Latin America, Remote

Bluelight Consulting 11-50 Internet Software & Services

Bluelight is hiring a DevOps/SRE/DevSecOps professional to operate secure, reliable AWS-based systems and delivery infrastructure for client software projects in a remote consultancy environment.

AWS CI/CD DevSecOps Docker K6 OpenTelemetry PostgreSQL Secrets Management Terraform
2 days, 10 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers