Staff Site Reliability Engineer

1 day, 11 hours ago
Full-time
Lead
DevOps and Infrastructure
Filevine

Filevine

Filevine is a top legal tech company revolutionizing legal work with AI-powered case management software, empowering law firms to streamline operations and enhance client services.

Specialized Consumer Services
251-1K
Founded 2015
$226M raised

Description

  • Define and execute the technical strategy for observability, alerting, platform infrastructure, and operational excellence.
  • Lead the evolution of reliable, scalable, secure, and efficient cloud platforms and distributed systems.
  • Champion SLIs, SLOs, error budgets, capacity planning, operational readiness, and automation across the service lifecycle.
  • Lead complex production incidents and drive post-incident learnings into permanent engineering improvements.
  • Build self-service platform capabilities that reduce toil, improve safety, and increase engineering velocity.
  • Mentor engineers and serve as a trusted technical authority for long-term reliability and platform direction.
  • Partner with engineering leadership and the Reliability Architect on major technical decisions and reliability strategy.
  • Shape on-call strategy, tooling, and culture to make production support more sustainable.
  • Drive the use of AI and machine learning in observability, anomaly detection, incident response, automated remediation, and resource optimization.

Requirements

  • 12+ years of experience in software engineering, infrastructure, platform engineering, or SRE.
  • 6+ years of experience in SRE.
  • 3+ years leading complex, cross-functional technical initiatives for distributed production systems.
  • Expert-level depth in observability and platform infrastructure.
  • Broad expertise in incident response, capacity planning, automation, and reliability engineering.
  • Advanced experience with a major container-orchestration platform, preferably Kubernetes.
  • Experience with an observability platform such as New Relic, Datadog, or equivalent.
  • Strong software engineering ability in Python, Go, Bash, or another general-purpose language.
  • Experience building production tooling, automation, or platform capabilities.
  • Experience in a regulated environment such as FedRAMP, CJIS, HIPAA, SOC 2, or PCI is strongly preferred.

Benefits

  • Medical, dental, and vision insurance for full-time employees.
  • Competitive and fair pay.
  • Maternity and paternity leave for full-time employees.
  • Short- and long-term disability coverage.
  • Opportunity to learn from a dedicated leadership team.
  • A dynamic, rapidly growing company environment.
  • Top-of-the-line company swag.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Software/Site Reliability Engineer - FedRAMP

Tenable 1K-5K Internet Software & Services

Tenable is hiring a Site Reliability Engineer to help scale and operate its cloud-based vulnerability management platform for private and U.S. Government cloud customers.

Agile AWS Azure Bash CI/CD Datadog Docker DynamoDB Elasticsearch GCP Go Gradle Groovy Helm Java Kafka Kotlin Kubernetes Microservices Node.js OpenSearch OpenTelemetry Python Splunk Terraform
9 hours, 18 minutes ago

Senior Site Reliability Engineer (SRE)

Branch 51-250 Professional Services

Branch is hiring a Senior Site Reliability Engineer to improve the reliability, scalability, performance, and observability of its fintech platform through automation and operational best practices.

Bash Docker GCP Go Gradle Grafana Java Kubernetes MySQL OpenTelemetry Prometheus Python Redis Spring Boot Terraform
9 hours, 48 minutes ago

Platform Engineering Manager

Prolific 51-250 Professional Services

Prolific is hiring a Platform Engineering Manager to lead its Cloud Platform and SRE teams, owning the technical foundation, reliability, and scalability of the infrastructure that supports its AI-focused human data platform.

Argo CD AWS Celery CircleCI Datadog DynamoDB Elasticsearch GCP GitHub Actions GitOps JavaScript Kubernetes MongoDB PostgreSQL Python Serverless Terraform TypeScript
2 days, 10 hours ago

Senior Site Reliability Engineer

Develocity is hiring founding Site Reliability Engineers to help build and operate its remote-first SaaS platform, ensuring reliability, performance, and availability for customer-facing and supporting services.

AWS Bash EC2 Grafana Java JUnit Kotlin Kubernetes Prometheus Python Spring Terraform
2 days, 11 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers