Megaport

Megaport

Megaport simplifies network connectivity with scalable bandwidth for cloud connections, metro ethernet, and Data Centre backhaul. Offering extensive coverage in APAC and expanding globally, Megaport empowers users to manage their networks through its u...

Diversified Telecommunication Services
251-1K
Founded 2013
$26M raised

Description

  • Improve production reliability and system resilience within an SRE team.
  • Champion engineering standards, DevOps practices, and industry best practices.
  • Communicate with internal teams and stakeholders throughout requirements analysis and demonstrations.
  • Investigate complex technical problems and develop effective solutions.
  • Write code, handle alerts, improve systems, and support team members.
  • Participate in on-call rotations, incident response, and blameless post-incident reviews.
  • Work across a broad and evolving technology environment.
  • Create and maintain runbooks and operational documentation.
  • Collaborate across time zones in an asynchronous, globally distributed organization.

Requirements

  • 5+ years of experience administering Linux systems and related production infrastructure.
  • Understanding of SRE concepts including SLIs, SLOs, SLAs, error budgets, blast radius, and blameless postmortems.
  • Strong focus on automation, reducing toil, and preventing recurring incidents.
  • Strong Kubernetes and ecosystem fundamentals.
  • Cloud infrastructure experience; AWS is strongly preferred and bare-metal experience is a bonus.
  • Strong Bash development skills, plus Python, Go, or a similar language.
  • Experience with infrastructure as code; Terraform is preferred.
  • Experience with CI/CD and version control; GitHub is preferred.
  • Experience with at least one of PostgreSQL, Cassandra, or ClickHouse.
  • Experience operating production observability systems for metrics, logs, and traces.
  • Strong troubleshooting skills and ownership of live production incident response.
  • Self-directed approach suited to an async, globally distributed team, with a commitment to continual professional development.

Benefits

  • Remote-first flexible work environment with coworking options.
  • Four weeks of paid annual leave, parental leave, birthday leave, and purchased annual leave options.
  • Wellness allowance and employee wellbeing initiatives.
  • Study and training allowance plus five days of paid study leave.
  • Modern workspaces for in-office collaboration.
  • Inclusive team environment with industry experts and developing talent.
  • Recognition programs including Legend and Kudos awards.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Site Reliability Engineer II

GoGuardian 251-1K Internet Software & Services

GoGuardian is seeking a Site Reliability Engineer II to build, maintain, and scale the cloud infrastructure, data services, and developer tooling that support reliable and secure K-12 learning products.

AWS CI/CD GitHub Actions Go JavaScript Jenkins Kubernetes Linux MongoDB OpenSearch Python Serverless Terraform TypeScript
1 day, 6 hours ago

Lead Site Reliability Engineer - Imunify Reliability Platform (remote work)

CloudLinux 51-250 IT Services

CloudLinux is seeking a Production Engineering/SRE specialist to establish observability, SLOs, alerting, and escalation practices for Imunify360’s distributed Linux security product, reducing detection of silent control failures from months to hours.

Ansible CI/CD ClickHouse GitLab CI Go Grafana Jenkins Kubernetes Linux OpenTelemetry Prometheus Python Rust WAF
1 day, 7 hours ago

Site Reliability Engineer, Tech Lead

Loadsmart 251-1K Air Freight & Logistics

Loadsmart is hiring a remote SRE Tech Lead in Brazil to build and operate its internal engineering platform, improve reliability, and enable safe, dependable applications across engineering teams.

Ansible AWS Bash Chef CI/CD Docker Kubernetes PostgreSQL Python Terraform
2 days, 7 hours ago

Senior Site Reliability Engineer (Performance and Scalability)

Digitalzone 251-1K Media

DigitalZone is hiring an SRE/platform engineer to build the scalability, observability, and resilience foundations that help engineering teams handle large campaign traffic spikes reliably.

AWS Go Laravel PHP PostgreSQL TypeScript
3 days, 7 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers