Tyk API Management

Tyk API Management

Tyk is a leading API Management Platform that enables interconnectivity between systems and devices through its fast, scalable, and open-source API Gateway, Analytics, Dev Portal, and Dashboard.

Internet Software & Services
51-250
Founded 2015
$40M raised

Description

  • Maintain Tyk Cloud availability and help define SLA/SLO/SI targets.
  • Identify reliability issues and work with the squad to resolve them.
  • Create and improve metrics and dashboards to monitor platform health.
  • Participate in the on-call rotation and serve as first-line incident management support.
  • Conduct post-incident analysis and help define response processes.
  • Automate common operational tasks and improve support workflows.
  • Document operational knowledge, SRE processes, and policies.
  • Support the expansion of the platform across multi-region and multi-cloud environments.
  • Recommend and implement ways to improve operational efficiency and reduce running costs without affecting service.
  • Assist with cloud penetration testing by coordinating with the provider and preparing technical details and environment setup.

Requirements

  • Experience launching and operating production-scale Kubernetes clusters.
  • Experience designing and operating infrastructure on AWS and other cloud providers.
  • Experience operating MongoDB or similar document databases.
  • Experience operating Redis or similar key-value storage clusters.
  • Experience administering Linux servers and maintaining distributed software.
  • Experience operating Prometheus, Grafana, and logging collection/analysis systems.
  • Strong collaboration skills and a proactive, energetic, innovative, change-oriented mindset.
  • Advanced knowledge of Kubernetes and containers, AWS/EKS, and Linux.
  • Proficient with Terraform and infrastructure as code, and Helm.
  • Familiarity with Go, monitoring tools such as Thanos, and networking concepts including subnets, routing, peering, load balancing, NAT, DNS, TCP/IP, HTTP, TLS, and UDP.
  • Availability to participate in the on-call rotation, including 16:00–4:00 UTC.
  • Nice to have: experience with GCP or Azure, bare metal infrastructure, API management, large-scale distributed storage, Rancher, CKA/CKAD/CKS certifications, or production software delivery in Go.

Benefits

  • Unlimited paid holiday.
  • Remote working from anywhere in the world.
  • Flexible working hours.
  • Employee share scheme.
  • Generous maternity and paternity leave.
  • Company retreats.
  • An inclusive, values-driven culture that emphasizes authenticity, respect, responsibility, independence, honesty, diversity, and inclusion.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

SRE Engineer Contractor

Beauty For All Industries (BFA) 51-200 information technology & services

IPSY is seeking a remote Site Reliability Engineer Contractor in Mexico or Colombia, covering the PST time zone, to improve the availability, resilience, and operational reliability of its beauty membership platform.

Amplitude AWS Bash CI/CD Contentful Datadog Grafana Microservices Netlify New Relic OpsGenie PagerDuty Prometheus Python Terraform
6 hours, 13 minutes ago

[Job 31894] AI ORCHESTRATOR (APP SRE)

CI&T 5K-10K Internet Software & Services

A CI&T busca um AI Orchestrator para governar a entrega ponta a ponta de um projeto de monitoramento SRE em jornadas de aplicativos de saúde, coordenando equipes e agentes de IA desde o upstream até a produção.

Agile
2 days, 7 hours ago

Senior Site Reliability Engineer (SRE)

UJET 251-1K Professional Services

UJET is seeking a Senior Site Reliability Engineer to build and scale its SRE function, improving reliability, reducing operational toil, and establishing production engineering best practices for its cloud-based contact center platform.

AWS Azure Go Java Kubernetes Microservices Python Terraform
3 days, 6 hours ago

[31893] AI ENGINEER SRE (APP)

CI&T 5K-10K Internet Software & Services

Engenheiro(a) SRE/Software Sênior na CI&T, atuando na rearquitetura e melhoria de performance de aplicativos de saúde, desde a definição da arquitetura até a entrega de soluções testadas para produção.

Angular AWS CI/CD Confluence Java JIRA Kanban Microservices Node.js Scrum Spring Boot
4 days, 6 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers