Tyk API Management

Tyk API Management

Tyk is a leading API Management Platform that enables interconnectivity between systems and devices through its fast, scalable, and open-source API Gateway, Analytics, Dev Portal, and Dashboard.

Internet Software & Services
51-250
Founded 2015
$40M raised

Description

  • Lead hands-on maintenance and optimization of the global cloud platform within defined SLAs, SLOs, and SLIs.
  • Collaborate with the SRE team to shape strategy and translate it into actionable technical plans.
  • Identify reliability issues, perform root cause analysis, and implement corrective solutions with the squad.
  • Lead performance tuning and fault-finding using OS and application metrics.
  • Design and implement automation for operational tasks and cloud operations workflows.
  • Develop monitoring, alerting, dashboards, and KPIs to improve platform visibility and response.
  • Participate in on-call rotation and support effective incident response, resolution, and postmortems.
  • Document operational findings, maintain runbooks, and drive continuous improvement across processes and practices.
  • Support multi-region and multi-cloud expansion with a focus on scalability and automation.
  • Engage with commercial teams on growth plans and translate them into technical SRE strategy.
  • Coordinate penetration testing and plan software upgrades to improve cloud services.

Requirements

  • Experience in an SRE role.
  • Strong knowledge of cloud technologies and SLA, SLO, and SLI management.
  • Experience with software design, automation, and root cause analysis.
  • Experience supporting production systems on-call with a customer-focused mindset.
  • Excellent communication and leadership skills.
  • Ability to analyze and improve operational processes and performance metrics.
  • Hands-on experience launching and operating production Kubernetes clusters.
  • Experience designing and operating infrastructure on AWS and other cloud providers.
  • Experience operating MongoDB or another document database, Redis or another key-value store, and Linux servers.
  • Experience with Prometheus, Grafana, and logging collection/analysis systems.
  • Advanced knowledge of Go, AWS/EKS, and Linux.
  • Proficient with Terraform and infrastructure as code, plus Helm.
  • Familiarity with monitoring tools such as Prometheus, Grafana, and Thanos.
  • Strong grasp of networking concepts and protocols such as DNS, TCP/IP, HTTP, TLS, UDP, subnets, routing, peering, load balancing, and NAT.
  • Ability to participate in the on-call rotation, including early-morning coverage from 4:00am to 16:00pm UTC.
  • Proactive, energetic, innovative, and change-oriented, with a desire to lead or mentor a team.

Benefits

  • Unlimited paid holidays.
  • Remote working from anywhere in the world.
  • Flexible working hours.
  • Employee share scheme.
  • Generous maternity and paternity leave.
  • Volunteering days.
  • Employee wellbeing platform.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

SRE Engineer Contractor

Beauty For All Industries (BFA) 51-200 information technology & services

IPSY is seeking a remote Site Reliability Engineer Contractor in Mexico or Colombia, covering the PST time zone, to improve the availability, resilience, and operational reliability of its beauty membership platform.

Amplitude AWS Bash CI/CD Contentful Datadog Grafana Microservices Netlify New Relic OpsGenie PagerDuty Prometheus Python Terraform
6 hours, 58 minutes ago

[Job 31894] AI ORCHESTRATOR (APP SRE)

CI&T 5K-10K Internet Software & Services

A CI&T busca um AI Orchestrator para governar a entrega ponta a ponta de um projeto de monitoramento SRE em jornadas de aplicativos de saúde, coordenando equipes e agentes de IA desde o upstream até a produção.

Agile
2 days, 7 hours ago

Senior Site Reliability Engineer (SRE)

UJET 251-1K Professional Services

UJET is seeking a Senior Site Reliability Engineer to build and scale its SRE function, improving reliability, reducing operational toil, and establishing production engineering best practices for its cloud-based contact center platform.

AWS Azure Go Java Kubernetes Microservices Python Terraform
3 days, 6 hours ago

[31893] AI ENGINEER SRE (APP)

CI&T 5K-10K Internet Software & Services

Engenheiro(a) SRE/Software Sênior na CI&T, atuando na rearquitetura e melhoria de performance de aplicativos de saúde, desde a definição da arquitetura até a entrega de soluções testadas para produção.

Angular AWS CI/CD Confluence Java JIRA Kanban Microservices Node.js Scrum Spring Boot
4 days, 7 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers