Site Reliability Engineering (SRE) Leader

1 week, 2 days ago
Full-time
Lead
DevOps and Infrastructure
PatSnap

PatSnap

PatSnap is a leading global patent and innovation database that uses deep learning algorithms to provide game-changing insights for innovation teams, revolutionizing collaboration and decision-making in the innovation lifecycle.

Internet Software & Services
251-1K
Founded 2007
$352M raised

Description

  • Build, lead, and develop the UK SRE team by establishing operational standards, best practices, and reliability goals.
  • Define and drive the operational strategy for the global SaaS platform.
  • Ensure high availability, stability, security, and performance of business-critical platforms and services.
  • Lead major incident management as the senior escalation point during critical production events.
  • Establish and monitor reliability metrics, including SLIs, SLOs, and operational KPIs.
  • Drive automation across infrastructure, deployments, monitoring, and operational workflows.
  • Champion the adoption of AI-powered operations to improve engineering productivity and operational excellence.
  • Partner with Engineering, Product, Security, and Infrastructure teams to improve architecture, scalability, and operational readiness.
  • Lead disaster recovery planning, operational resilience initiatives, and platform risk management.
  • Evaluate emerging cloud, AI, and platform technologies to support continuous improvement.

Requirements

  • Bachelor’s degree in Computer Science or a related field.
  • At least 8 years of experience in DevOps, SRE, or infrastructure operations.
  • Proven experience leading technical teams and managing production environments at scale.
  • Strong expertise in cloud platforms, with AWS preferred.
  • Experience with Kubernetes, Docker, CI/CD pipelines, Infrastructure as Code, and observability platforms.
  • Deep understanding of distributed systems, high-availability architectures, and large-scale SaaS environments.
  • Experience driving automation and operational excellence initiatives.
  • Hands-on experience using AI tools such as ChatGPT, Claude, GitHub Copilot, Codex, or similar technologies.
  • Strong problem-solving, leadership, communication, and stakeholder management skills.
  • Fluent in English; Mandarin is highly desirable for collaboration across regions.

Benefits

  • Work for a pre-IPO, global company with a $1B valuation and strong growth ahead.
  • Join a company trusted by more than 15,000 organisations worldwide.
  • Opportunity to shape an AI-powered SRE function and influence engineering strategy.
  • Collaborate with teams across the UK, Singapore, China, and other global locations.
  • Be part of a diverse team with offices in Singapore, Toronto, London, and Shanghai, plus remote teams in the US.
  • Equal opportunity employer with accommodations available during the interview process.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Director, Engineering - Infrastructure

Federato 11-50 Insurance

Federato is hiring a Director of Infrastructure to lead the platform, SRE, and Security functions that keep its AI-native insurance workflow system scalable, reliable, secure, and cost-efficient.

1 day, 4 hours ago

Technical Support Engineer (GPU Clusters) - US Weekends

Together 1-10 IT Services

Together AI is hiring a Technical Support Engineer to support customers building and operating AI training, fine-tuning, and inference systems on Kubernetes GPU infrastructure.

Ansible Kubernetes Machine Learning
2 days, 4 hours ago

Staff Site Reliability Engineer

BeyondTrust 1K-5K Professional Services

BeyondTrust is hiring a Staff Site Reliability Engineer to lead the evolution of its Password Safe platform, infrastructure, and deployment ecosystem across cloud and on-premises environments.

Ansible AWS Azure C# CI/CD Datadog DevSecOps Docker GitOps Go Java Kubernetes Linux Microservices OpenTelemetry Secrets Management Terraform
2 days, 4 hours ago

Senior Site Reliability Engineer

Alpaca 51-250 Capital Markets

Alpaca is hiring a Site Reliability Engineer to keep its brokerage platform reliable, observable, and operable across cloud infrastructure, Kubernetes, and PostgreSQL on the trading-critical path.

DNS GitOps Go Kafka Kubernetes Linux Load Balancing PostgreSQL Python RabbitMQ Secrets Management TLS
3 days, 3 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers