Airalo

Airalo

Airalo is the world's first eSIM store offering travelers access to eSIMs in 200+ countries & regions at affordable prices. With Airalo, travelers can manage their eSIMs, top up on the go, and enjoy pain-free connectivity while traveling. Say goodbye t...

Airlines
51-250
Founded 2019
$67M raised

Description

  • Lead the design of scalable, fault-tolerant, self-healing systems in a multi-region AWS environment.
  • Define and track SLOs and SLIs to guide architectural decisions and error budget policies.
  • Conduct blameless post-incident reviews and implement long-term preventive measures.
  • Identify manual operational work and build internal tools and automation to eliminate it.
  • Develop and maintain automated runbooks and playbooks for operations and incident response.
  • Improve observability by turning monitoring data into proactive, actionable insights.
  • Proactively identify and mitigate operational risks through chaos engineering and architecture reviews.
  • Partner with software engineers to design for reliability, scalability, and maintainability early in the SDLC.
  • Continuously evaluate and optimize system performance, capacity, and cost efficiency.
  • Refine the on-call experience to reduce alert fatigue, improve MTTR, and maintain sustainable rotations.

Requirements

  • Bachelor’s degree in Computer Engineering or a similar discipline.
  • 5+ years of experience as a Site Reliability Engineer or in a similar role.
  • 3+ years of experience with AWS services, including strong knowledge of container orchestration.
  • 2+ years of Kubernetes experience.
  • Deep understanding of observability principles and tools such as Prometheus, Datadog, and OpenTelemetry.
  • Experience with leading incident management and complex postmortem analysis.
  • Experience and interest in infrastructure as code, especially Terraform.
  • Experience with chaos engineering and other resilience-testing techniques.
  • Experience with CI/CD tools such as GitHub Actions for automated delivery.
  • Proficiency in at least one programming language such as Python, Go, or Java for automation and internal tooling.
  • Event-driven architecture experience with tools such as SNS and SQS.
  • Ability to work independently and collaboratively in a fast-paced environment.
  • Good communication skills and fluency in English.
  • Prior experience with Scrum or other agile methods (preferred).
  • Certification such as AWS Certified DevOps Engineer or Certified Kubernetes Administrator (preferred).
  • Prior experience with Telco core networks, low-latency networking, telecommunications, eSIM, or GSMA technologies (preferred).
  • Experience with AI-driven SRE tools for anomaly detection and improvement (preferred).
  • Contributions to open-source SRE projects or communities (preferred).

Benefits

  • Remote work.
  • Generous PTO.
  • Wellness allowance.
  • Learning allowance.
  • Annual Airalo Away retreat.
  • Paid on-call standby fees plus overtime pay.
  • Guaranteed rest periods and flexible hours after night incidents.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

Counterpart Health 51-200 hospital & health care

Counterpart Health is hiring a Senior Site Reliability and Infrastructure Engineer to support and evolve the technology platform behind its primary care tool and maintain reliable infrastructure for domestic and international workloads.

AWS Azure CI/CD Containerd DNS Docker GCP Go gRPC Helm Kubernetes Linux Load Balancing Prometheus Python Shell Scripting TCP/IP
16 hours, 1 minute ago

Senior Test Platform & Reliability Engineer - Star Trek Fleet Command

Scopely 1K-5K Internet Software & Services

Scopely is hiring a Senior Test Platform & Reliability Engineer in Ireland to build validation, reliability, and developer enablement platforms for Star Trek Fleet Command’s large-scale live-service backend systems.

AWS Bash CI/CD Docker GitLab Go Python Terraform
16 hours, 16 minutes ago

Senior Software Engineer - Databases, SRE | Canada | Remote

Grafana 1K-5K IT Services

Grafana Labs is hiring a Senior Software Engineer for its remote SRE team to improve reliability and operability of Grafana Cloud database services for high-SLA customers across AWS, GCP, and Azure.

AWS Azure GCP Go Helm Java Kubernetes Linux Microservices Python Terraform
1 day, 15 hours ago

Senior Site Reliability Engineer

Semios 51-250 Food Products

Semios Group is hiring a Senior Site Reliability Engineer to help scale, secure, and improve the reliability of its global agricultural technology platform.

AWS Azure Bash Buildkite CI/CD Datadog Docker Envoy GCP Git GitHub GitHub Actions GitLab Go Jenkins Kubernetes Linux NATS New Relic Prometheus Python Ruby Splunk Terraform
1 day, 16 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers