Senior Site Reliability engineer

1 month, 1 week ago
Full-time
Lead
DevOps and Infrastructure
Filevine

Filevine

Filevine is a top legal tech company revolutionizing legal work with AI-powered case management software, empowering law firms to streamline operations and enhance client services.

Specialized Consumer Services
251-1K
Founded 2015
$226M raised

Description

  • Design and improve monitoring, logging, distributed tracing, dashboards, alerting, SLIs, and SLOs for production visibility.
  • Build and maintain automation, internal tools, and CI/CD systems that reduce toil and support reliable deployments at scale.
  • Improve systems for building, testing, deploying, and operating Filevine products while proactively addressing reliability, performance, scalability, and security risks.
  • Own complex production incidents from detection and triage through communication, resolution, and follow-up.
  • Turn incident learnings into corrective actions, stronger runbooks, and improved operational practices.
  • Lead significant technical initiatives from problem definition and design through implementation and adoption.
  • Coordinate work across engineers and teams, communicate tradeoffs and risks, and ensure delivery of intended results.
  • Mentor other Site Reliability Engineers through design reviews, incident follow-ups, and paired problem-solving.
  • Participate in the shared on-call rotation and support capacity planning, operational readiness, resilience, and recovery improvements.
  • Apply AI and machine learning to operational signals to identify patterns, forecast reliability and capacity risks, and improve system operations.

Requirements

  • 8+ years of hands-on experience in software engineering, cloud infrastructure, platform engineering, DevOps, or related technical roles.
  • At least 5 years of experience in a Site Reliability Engineering or reliability-focused role.
  • Strong knowledge of distributed systems and hands-on experience operating Kubernetes workloads.
  • Experience with cloud infrastructure in AWS or a comparable platform.
  • Proficiency in Infrastructure as Code, monitoring, logging, alerting, distributed tracing, SLIs, and SLOs.
  • Strong proficiency with Python, Go, Bash, or a similar language.
  • Experience building and maintaining production tooling, automation, CI/CD pipelines, and deployment systems.
  • Demonstrated ability to lead troubleshooting, incident response, root cause analysis, and long-term reliability improvements.
  • Proven ability to mentor SREs and communicate clearly with technical and business stakeholders.
  • Experience applying AI and machine learning to operational data and engineering workflows, with appropriate safeguards.

Benefits

  • Medical, dental, and vision insurance for full-time employees.
  • Competitive and fair pay.
  • Maternity and paternity leave for full-time employees.
  • Short-term and long-term disability coverage.
  • Opportunity to learn from a dedicated leadership team.
  • A dynamic, rapidly growing company environment.
  • Top-of-the-line company swag.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Service Reliability Engineer

Thoughtworks 10K-50K Professional Services

Thoughtworks is seeking a Senior Site Reliability Engineer to help clients improve infrastructure reliability, resilience, observability, and operational performance through automation and continuous improvement.

AWS Azure CI/CD Datadog ELK Stack GitOps Grafana Kubernetes Microservices New Relic Nomad Python REST API
10 hours, 53 minutes ago

Cloud Site Reliability Engineer

Cadwell 51-250 Health Care Providers & Services

Cadwell is seeking a Cloud Site Reliability Engineer to operate and improve AWS infrastructure supporting healthcare customers and ensure reliable, secure, and compliant hosted neurodiagnostic software environments.

AWS Bash CI/CD Encryption HIPAA JavaScript JSON Python Terraform TypeScript YAML
1 day, 10 hours ago

Site Reliability Engineer

GiveCampus 51-250 Internet Software & Services

GiveCampus is seeking a hands-on Site Reliability Engineer to strengthen the reliability, performance, observability, and operational maturity of its AWS-based fundraising platform in a remote-first U.S. role.

AWS CI/CD CircleCI Datadog GitHub Actions Kubernetes Linux New Relic OpenSearch PostgreSQL Redis Ruby Ruby on Rails Terraform
1 day, 10 hours ago

DevOps / SRE / DevSecOps Engineer (AWS) - Latin America, Remote

Bluelight Consulting 11-50 Internet Software & Services

Bluelight is hiring a DevOps/SRE/DevSecOps professional to operate secure, reliable AWS-based systems and delivery infrastructure for client software projects in a remote consultancy environment.

AWS CI/CD DevSecOps Docker K6 OpenTelemetry PostgreSQL Secrets Management Terraform
2 days, 10 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers