Senior Site Reliability engineer

13 hours, 58 minutes ago
Full-time
Lead
DevOps and Infrastructure
Filevine

Filevine

Filevine is a top legal tech company revolutionizing legal work with AI-powered case management software, empowering law firms to streamline operations and enhance client services.

Specialized Consumer Services
251-1K
Founded 2015
$226M raised

Description

  • Design and improve monitoring, logging, distributed tracing, dashboards, alerting, SLIs, and SLOs for production visibility.
  • Build and maintain automation, internal tools, and CI/CD systems that reduce toil and support reliable deployments at scale.
  • Improve systems for building, testing, deploying, and operating Filevine products while proactively addressing reliability, performance, scalability, and security risks.
  • Own complex production incidents from detection and triage through communication, resolution, and follow-up.
  • Turn incident learnings into corrective actions, stronger runbooks, and improved operational practices.
  • Lead significant technical initiatives from problem definition and design through implementation and adoption.
  • Coordinate work across engineers and teams, communicate tradeoffs and risks, and ensure delivery of intended results.
  • Mentor other Site Reliability Engineers through design reviews, incident follow-ups, and paired problem-solving.
  • Participate in the shared on-call rotation and support capacity planning, operational readiness, resilience, and recovery improvements.
  • Apply AI and machine learning to operational signals to identify patterns, forecast reliability and capacity risks, and improve system operations.

Requirements

  • 8+ years of hands-on experience in software engineering, cloud infrastructure, platform engineering, DevOps, or related technical roles.
  • At least 5 years of experience in a Site Reliability Engineering or reliability-focused role.
  • Strong knowledge of distributed systems and hands-on experience operating Kubernetes workloads.
  • Experience with cloud infrastructure in AWS or a comparable platform.
  • Proficiency in Infrastructure as Code, monitoring, logging, alerting, distributed tracing, SLIs, and SLOs.
  • Strong proficiency with Python, Go, Bash, or a similar language.
  • Experience building and maintaining production tooling, automation, CI/CD pipelines, and deployment systems.
  • Demonstrated ability to lead troubleshooting, incident response, root cause analysis, and long-term reliability improvements.
  • Proven ability to mentor SREs and communicate clearly with technical and business stakeholders.
  • Experience applying AI and machine learning to operational data and engineering workflows, with appropriate safeguards.

Benefits

  • Medical, dental, and vision insurance for full-time employees.
  • Competitive and fair pay.
  • Maternity and paternity leave for full-time employees.
  • Short-term and long-term disability coverage.
  • Opportunity to learn from a dedicated leadership team.
  • A dynamic, rapidly growing company environment.
  • Top-of-the-line company swag.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

PandaDoc 251-1K Internet Software & Services

PandaDoc is hiring a Site Reliability Engineer to keep its document workflow platform highly available and resilient while supporting production operations and reliability improvements.

Agile AWS Django Grafana Java Kafka Kubernetes NATS PostgreSQL Python RabbitMQ Spring Boot
12 hours, 58 minutes ago

Senior Site Reliability Engineer

PandaDoc 251-1K Internet Software & Services

PandaDoc is hiring a Site Reliability Engineer to help keep its production document workflow platform reliable, resilient, and low-downtime for customers.

AWS Django Grafana Java Kafka Kubernetes NATS PostgreSQL Python RabbitMQ Spring Boot
13 hours, 12 minutes ago

Senior Site Reliability Engineer

PandaDoc 251-1K Internet Software & Services

PandaDoc is hiring a Site Reliability Engineer to keep its document workflow platform highly available, resilient, and efficient for customers.

AWS Django Grafana Java Kafka Kubernetes NATS PostgreSQL Python RabbitMQ Spring Boot
13 hours, 12 minutes ago

Director, Engineering - Infrastructure

Federato 11-50 Insurance

Federato is hiring a Director of Infrastructure to lead the platform, SRE, and Security functions that keep its AI-native insurance workflow system scalable, reliable, secure, and cost-efficient.

2 days, 12 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers