Tyk API Management

Tyk API Management

Tyk is a leading API Management Platform that enables interconnectivity between systems and devices through its fast, scalable, and open-source API Gateway, Analytics, Dev Portal, and Dashboard.

Internet Software & Services
51-250
Founded 2015
$40M raised

Description

  • Maintain Tyk Cloud availability and help define SLA/SLO/SI targets.
  • Identify reliability issues and work with the squad to resolve them.
  • Create and improve metrics and dashboards to monitor platform health.
  • Participate in the on-call rotation and serve as first-line incident management support.
  • Conduct post-incident analysis and help define response processes.
  • Automate common operational tasks and improve support workflows.
  • Document operational knowledge, SRE processes, and policies.
  • Support the expansion of the platform across multi-region and multi-cloud environments.
  • Recommend and implement ways to improve operational efficiency and reduce running costs without affecting service.
  • Assist with cloud penetration testing by coordinating with the provider and preparing technical details and environment setup.

Requirements

  • Experience launching and operating production-scale Kubernetes clusters.
  • Experience designing and operating infrastructure on AWS and other cloud providers.
  • Experience operating MongoDB or similar document databases.
  • Experience operating Redis or similar key-value storage clusters.
  • Experience administering Linux servers and maintaining distributed software.
  • Experience operating Prometheus, Grafana, and logging collection/analysis systems.
  • Strong collaboration skills and a proactive, energetic, innovative, change-oriented mindset.
  • Advanced knowledge of Kubernetes and containers, AWS/EKS, and Linux.
  • Proficient with Terraform and infrastructure as code, and Helm.
  • Familiarity with Go, monitoring tools such as Thanos, and networking concepts including subnets, routing, peering, load balancing, NAT, DNS, TCP/IP, HTTP, TLS, and UDP.
  • Availability to participate in the on-call rotation, including 16:00–4:00 UTC.
  • Nice to have: experience with GCP or Azure, bare metal infrastructure, API management, large-scale distributed storage, Rancher, CKA/CKAD/CKS certifications, or production software delivery in Go.

Benefits

  • Unlimited paid holiday.
  • Remote working from anywhere in the world.
  • Flexible working hours.
  • Employee share scheme.
  • Generous maternity and paternity leave.
  • Company retreats.
  • An inclusive, values-driven culture that emphasizes authenticity, respect, responsibility, independence, honesty, diversity, and inclusion.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Staff Site Reliability Engineer-Observability

GoDaddy 5001-10000 Technology, Information and Internet

GoDaddy is hiring a remote Reliability Engineer in India to operate and modernize the monitoring, compliance, infrastructure, and incident-response systems supporting its global Domains platform.

Ansible Argo CD AWS AWS CDK CI/CD Cybersecurity Elasticsearch Git Go Gradle Grafana Jenkins Kubernetes Maven MySQL PostgreSQL Prometheus Python Ruby SQL Terraform
15 hours, 3 minutes ago

Site Reliability Engineer - AWS and Azure

Jalasoft 1K-5K Internet Software & Services

Jalasoft is hiring a Site Reliability Engineer to build, maintain, and improve reliable, scalable, and secure cloud infrastructure across Windows and Linux environments.

Ansible AWS Azure Bash CI/CD Docker GitHub GitHub Actions Grafana Kubernetes PowerShell Prometheus Python TeamCity Terraform
1 day, 15 hours ago

[Job-31445] Sênior Software Engineer | SRE & Software Architecture, Brazil

CI&T 5K-10K Internet Software & Services

A CI&T busca uma pessoa especialista em observabilidade para atuar como facilitadora técnica junto aos times de desenvolvimento, apoiando a confiabilidade, a performance e a evolução das aplicações.

Agile Angular Azure CI/CD Datadog Docker Git GitFlow GitHub Actions Grafana Java Kanban Kubernetes OpenShift OpenTelemetry Prometheus Scrum WAF
2 days, 15 hours ago

Senior Site Reliability Engineer

Sports Academy Education Services

Texas Sports Academy is seeking a part-time Senior Site Reliability Engineer consultant to audit, improve, and scale the infrastructure supporting its AI-first K–12 school.

AWS CI/CD Datadog Grafana Prometheus
2 days, 15 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers