Site Reliability Engineering (SRE) Leader

23 hours, 36 minutes ago
Full-time
Lead
DevOps and Infrastructure
PatSnap

PatSnap

PatSnap is a leading global patent and innovation database that uses deep learning algorithms to provide game-changing insights for innovation teams, revolutionizing collaboration and decision-making in the innovation lifecycle.

Internet Software & Services
251-1K
Founded 2007
$352M raised

Description

  • Build, lead, and develop the UK SRE team by establishing operational standards, best practices, and reliability goals.
  • Define and drive the operational strategy for the global SaaS platform.
  • Ensure high availability, stability, security, and performance of business-critical platforms and services.
  • Lead major incident management as the senior escalation point during critical production events.
  • Establish and monitor reliability metrics, including SLIs, SLOs, and operational KPIs.
  • Drive automation across infrastructure, deployments, monitoring, and operational workflows.
  • Champion the adoption of AI-powered operations to improve engineering productivity and operational excellence.
  • Partner with Engineering, Product, Security, and Infrastructure teams to improve architecture, scalability, and operational readiness.
  • Lead disaster recovery planning, operational resilience initiatives, and platform risk management.
  • Evaluate emerging cloud, AI, and platform technologies to support continuous improvement.

Requirements

  • Bachelor’s degree in Computer Science or a related field.
  • At least 8 years of experience in DevOps, SRE, or infrastructure operations.
  • Proven experience leading technical teams and managing production environments at scale.
  • Strong expertise in cloud platforms, with AWS preferred.
  • Experience with Kubernetes, Docker, CI/CD pipelines, Infrastructure as Code, and observability platforms.
  • Deep understanding of distributed systems, high-availability architectures, and large-scale SaaS environments.
  • Experience driving automation and operational excellence initiatives.
  • Hands-on experience using AI tools such as ChatGPT, Claude, GitHub Copilot, Codex, or similar technologies.
  • Strong problem-solving, leadership, communication, and stakeholder management skills.
  • Fluent in English; Mandarin is highly desirable for collaboration across regions.

Benefits

  • Work for a pre-IPO, global company with a $1B valuation and strong growth ahead.
  • Join a company trusted by more than 15,000 organisations worldwide.
  • Opportunity to shape an AI-powered SRE function and influence engineering strategy.
  • Collaborate with teams across the UK, Singapore, China, and other global locations.
  • Be part of a diverse team with offices in Singapore, Toronto, London, and Shanghai, plus remote teams in the US.
  • Equal opportunity employer with accommodations available during the interview process.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Sr. Site Reliability Engineer, tvScientific

Pinterest 5K-10K Internet Software & Services

Pinterest is hiring a Senior Site Reliability Engineer to operate and improve tvScientific’s cloud-native CTV advertising platform on AWS, Kubernetes, and GitOps workflows.

Argo CD AWS Bash CI/CD GCP GitHub Actions GitOps Helm Kubernetes Linux Python Secrets Management Terraform
22 hours, 36 minutes ago

SITE RELIABILITY ENGINEER III

Harford County Public Library 51-250 Diversified Consumer Services

Site Reliability Engineer na Stone, atuando no time de Foundation Platform para fortalecer a plataforma interna de tecnologia com foco em observabilidade, automação e estabilidade dos sistemas.

Ansible Argo CD AWS Azure Datadog Docker GCP GitHub Actions Go Grafana Kubernetes Linux Node.js OpenTelemetry Prometheus Python Splunk Terraform
1 day, 22 hours ago

Sr. Site Reliability Engineer (Starlink)

SpaceX 10K-50K Aerospace & Defense

SpaceX is hiring a Sr. Site Reliability Engineer for Starlink to improve the reliability, scalability, and performance of the systems supporting its satellite internet service.

Apache Spark C# CI/CD Flink Git Go HDFS Java Kafka Kubernetes Linux Python Scala
1 day, 22 hours ago

Head of Platform Engineering

dLocal 251-1K Diversified Financial Services

dLocal is seeking a senior leader to own its engineering platform, reliability posture, and AI-assisted development transformation across a global payments business serving emerging markets.

CI/CD Microservices
1 day, 23 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers