Senior Site Reliability engineer

3 weeks, 2 days ago
Full-time
Lead
DevOps and Infrastructure
Filevine

Filevine

Filevine is a top legal tech company revolutionizing legal work with AI-powered case management software, empowering law firms to streamline operations and enhance client services.

Specialized Consumer Services
251-1K
Founded 2015
$226M raised

Description

  • Design and improve monitoring, logging, distributed tracing, dashboards, alerting, SLIs, and SLOs for production visibility.
  • Build and maintain automation, internal tools, and CI/CD systems that reduce toil and support reliable deployments at scale.
  • Improve systems for building, testing, deploying, and operating Filevine products while proactively addressing reliability, performance, scalability, and security risks.
  • Own complex production incidents from detection and triage through communication, resolution, and follow-up.
  • Turn incident learnings into corrective actions, stronger runbooks, and improved operational practices.
  • Lead significant technical initiatives from problem definition and design through implementation and adoption.
  • Coordinate work across engineers and teams, communicate tradeoffs and risks, and ensure delivery of intended results.
  • Mentor other Site Reliability Engineers through design reviews, incident follow-ups, and paired problem-solving.
  • Participate in the shared on-call rotation and support capacity planning, operational readiness, resilience, and recovery improvements.
  • Apply AI and machine learning to operational signals to identify patterns, forecast reliability and capacity risks, and improve system operations.

Requirements

  • 8+ years of hands-on experience in software engineering, cloud infrastructure, platform engineering, DevOps, or related technical roles.
  • At least 5 years of experience in a Site Reliability Engineering or reliability-focused role.
  • Strong knowledge of distributed systems and hands-on experience operating Kubernetes workloads.
  • Experience with cloud infrastructure in AWS or a comparable platform.
  • Proficiency in Infrastructure as Code, monitoring, logging, alerting, distributed tracing, SLIs, and SLOs.
  • Strong proficiency with Python, Go, Bash, or a similar language.
  • Experience building and maintaining production tooling, automation, CI/CD pipelines, and deployment systems.
  • Demonstrated ability to lead troubleshooting, incident response, root cause analysis, and long-term reliability improvements.
  • Proven ability to mentor SREs and communicate clearly with technical and business stakeholders.
  • Experience applying AI and machine learning to operational data and engineering workflows, with appropriate safeguards.

Benefits

  • Medical, dental, and vision insurance for full-time employees.
  • Competitive and fair pay.
  • Maternity and paternity leave for full-time employees.
  • Short-term and long-term disability coverage.
  • Opportunity to learn from a dedicated leadership team.
  • A dynamic, rapidly growing company environment.
  • Top-of-the-line company swag.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Specialist II

Everbridge 1K-5K Internet Software & Services

Everbridge is hiring a Senior Site Reliability Engineer II to build resilient platforms and improve the reliability, scalability, and operational readiness of systems supporting critical-event communications.

CI/CD Kubernetes Linux
9 hours, 15 minutes ago

Site Reliability Engineer (SRE)

Rocket.net 11-50 IT Services

Rocket.net is seeking a Site Reliability Engineer to maintain the reliability and performance of its hosting platform while resolving complex infrastructure issues and providing advanced support to customers.

Apache Bash CDN Cloudflare Datadog DNS Linux MariaDB MySQL Nginx Redis SSH WAF WordPress
9 hours, 15 minutes ago

Senior Site Reliability Engineer

GoDaddy 5001-10000 Technology, Information and Internet

GoDaddy is hiring a Senior Site Reliability Engineer to support the Commerce ecosystem remotely by improving the reliability, scalability, security, and operation of business-critical production platforms.

Ansible AWS AWS CDK CI/CD CloudFormation Go Kubernetes Linux Pulumi Python SaltStack Terraform TypeScript
1 day, 8 hours ago

Senior Site Reliability Engineer

PENN Entertainment 10K-50K Hotels, Restaurants & Leisure

PENN Entertainment is seeking a Senior Site Reliability Engineer to build and operate the cloud infrastructure supporting large-scale sports betting and media platforms across regulated jurisdictions.

Argo CD AWS Bash CI/CD Datadog GCP GitHub Actions GitOps Go Helm Kubernetes Linux PostgreSQL Python Shell Scripting Terraform
2 days, 8 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers