Airalo

Airalo

Airalo is the world's first eSIM store offering travelers access to eSIMs in 200+ countries & regions at affordable prices. With Airalo, travelers can manage their eSIMs, top up on the go, and enjoy pain-free connectivity while traveling. Say goodbye t...

Airlines
51-250
Founded 2019
$67M raised

Description

  • Lead the design of scalable, fault-tolerant, self-healing systems in a multi-region AWS environment.
  • Define and track SLOs and SLIs to guide architectural decisions and error budget policies.
  • Conduct blameless post-incident reviews and implement long-term preventive measures.
  • Identify manual operational work and build internal tools and automation to eliminate it.
  • Develop and maintain automated runbooks and playbooks for operations and incident response.
  • Improve observability by turning monitoring data into proactive, actionable insights.
  • Proactively identify and mitigate operational risks through chaos engineering and architecture reviews.
  • Partner with software engineers to design for reliability, scalability, and maintainability early in the SDLC.
  • Continuously evaluate and optimize system performance, capacity, and cost efficiency.
  • Refine the on-call experience to reduce alert fatigue, improve MTTR, and maintain sustainable rotations.

Requirements

  • Bachelor’s degree in Computer Engineering or a similar discipline.
  • 5+ years of experience as a Site Reliability Engineer or in a similar role.
  • 3+ years of experience with AWS services, including strong knowledge of container orchestration.
  • 2+ years of Kubernetes experience.
  • Deep understanding of observability principles and tools such as Prometheus, Datadog, and OpenTelemetry.
  • Experience with leading incident management and complex postmortem analysis.
  • Experience and interest in infrastructure as code, especially Terraform.
  • Experience with chaos engineering and other resilience-testing techniques.
  • Experience with CI/CD tools such as GitHub Actions for automated delivery.
  • Proficiency in at least one programming language such as Python, Go, or Java for automation and internal tooling.
  • Event-driven architecture experience with tools such as SNS and SQS.
  • Ability to work independently and collaboratively in a fast-paced environment.
  • Good communication skills and fluency in English.
  • Prior experience with Scrum or other agile methods (preferred).
  • Certification such as AWS Certified DevOps Engineer or Certified Kubernetes Administrator (preferred).
  • Prior experience with Telco core networks, low-latency networking, telecommunications, eSIM, or GSMA technologies (preferred).
  • Experience with AI-driven SRE tools for anomaly detection and improvement (preferred).
  • Contributions to open-source SRE projects or communities (preferred).

Benefits

  • Remote work.
  • Generous PTO.
  • Wellness allowance.
  • Learning allowance.
  • Annual Airalo Away retreat.
  • Paid on-call standby fees plus overtime pay.
  • Guaranteed rest periods and flexible hours after night incidents.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

PENN Entertainment 10K-50K Hotels, Restaurants & Leisure

PENN Entertainment is seeking a Senior Site Reliability Engineer to build and operate the cloud infrastructure supporting large-scale sports betting and media platforms across regulated jurisdictions.

Argo CD AWS Bash CI/CD Datadog GCP GitHub Actions GitOps Go Helm Kubernetes Linux PostgreSQL Python Shell Scripting Terraform
16 hours, 54 minutes ago

Staff Site Reliability Engineer, Ads

Reddit 1K-5K Internet Software & Services

Reddit is hiring a Staff Site Reliability Engineer to provide technical leadership for reliability, scalability, and operational excellence across its advertising infrastructure and revenue-critical systems.

Apache Spark ClickHouse GCP Go Kafka Kubernetes
17 hours, 24 minutes ago

Sr Lead Network Reliability Engineer

Coupa Software 1K-5K Internet Software & Services

Coupa is hiring a Sr. Lead Network Development Engineer to scale and operate its global SaaS platform’s cloud networking infrastructure through automation, reliability engineering, and technical leadership.

Ansible AWS Azure Chef DNS Fortinet Go Java Kubernetes Linux Python Ruby TCP/IP Terraform TLS
1 day, 17 hours ago

Senior Monitoring/Observability Architect

Makpar 51-250 Internet Software & Services

Makpar is seeking a Senior Monitoring/Observability Architect to lead enterprise monitoring strategy, architecture, and implementation guidance for a large federal government program.

Datadog Splunk
1 day, 17 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers