Tyk API Management

Tyk API Management

Tyk is a leading API Management Platform that enables interconnectivity between systems and devices through its fast, scalable, and open-source API Gateway, Analytics, Dev Portal, and Dashboard.

Internet Software & Services
51-250
Founded 2015
$40M raised

Description

  • Maintain Tyk Cloud availability and help define SLA/SLO/SI targets.
  • Identify reliability issues and work with the squad to resolve them.
  • Create and improve metrics and dashboards to monitor platform health.
  • Participate in the on-call rotation and serve as first-line incident management support.
  • Conduct post-incident analysis and help define response processes.
  • Automate common operational tasks and improve support workflows.
  • Document operational knowledge, SRE processes, and policies.
  • Support the expansion of the platform across multi-region and multi-cloud environments.
  • Recommend and implement ways to improve operational efficiency and reduce running costs without affecting service.
  • Assist with cloud penetration testing by coordinating with the provider and preparing technical details and environment setup.

Requirements

  • Experience launching and operating production-scale Kubernetes clusters.
  • Experience designing and operating infrastructure on AWS and other cloud providers.
  • Experience operating MongoDB or similar document databases.
  • Experience operating Redis or similar key-value storage clusters.
  • Experience administering Linux servers and maintaining distributed software.
  • Experience operating Prometheus, Grafana, and logging collection/analysis systems.
  • Strong collaboration skills and a proactive, energetic, innovative, change-oriented mindset.
  • Advanced knowledge of Kubernetes and containers, AWS/EKS, and Linux.
  • Proficient with Terraform and infrastructure as code, and Helm.
  • Familiarity with Go, monitoring tools such as Thanos, and networking concepts including subnets, routing, peering, load balancing, NAT, DNS, TCP/IP, HTTP, TLS, and UDP.
  • Availability to participate in the on-call rotation, including 16:00–4:00 UTC.
  • Nice to have: experience with GCP or Azure, bare metal infrastructure, API management, large-scale distributed storage, Rancher, CKA/CKAD/CKS certifications, or production software delivery in Go.

Benefits

  • Unlimited paid holiday.
  • Remote working from anywhere in the world.
  • Flexible working hours.
  • Employee share scheme.
  • Generous maternity and paternity leave.
  • Company retreats.
  • An inclusive, values-driven culture that emphasizes authenticity, respect, responsibility, independence, honesty, diversity, and inclusion.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Manager Site Reliability Engineer

GoDaddy 5001-10000 Technology, Information and Internet

GoDaddy is hiring a Senior Manager of Site Reliability Engineering in India to lead its SRE Center of Excellence and build a global, platform-focused operations capability that improves reliability, developer productivity, and scale.

AWS CI/CD Datadog GitHub Actions Go Grafana Kafka Kubernetes Microservices PostgreSQL Prometheus Python Terraform
13 hours, 17 minutes ago

Vice President, Global Production Operations & Reliability

Everbridge 1K-5K Internet Software & Services

Everbridge is hiring a Vice President, Global Production Operations & Reliability to lead the company’s global production operations for its cloud-native SaaS platform and drive reliability, scalability, security, and operational excellence.

AWS CI/CD Kubernetes
1 day, 12 hours ago

DevOps Engineer - SRE Observability

Lingaro 5K-10K IT Services

An infrastructure-focused role at Lingaro responsible for monitoring, automating, and designing cloud systems within an Azure-based environment.

Azure Azure Pipelines CI/CD Docker GitHub GitHub Actions Grafana Kubernetes MySQL PostgreSQL Prometheus SQL Terraform
1 day, 12 hours ago

Staff Field Reliability Engineer

Honeycomb.io 51-250 Internet Software & Services

Honeycomb is hiring a Field Reliability Engineer to lead complex customer escalations, managed infrastructure operations, and observability strategy for its cloud-based platform.

Ansible AWS Chef EC2 Go Helm Honeycomb Java Kubernetes .NET Node.js OpenTelemetry Python Serverless Terraform TypeScript
1 day, 12 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers