Megaport

Megaport

Megaport simplifies network connectivity with scalable bandwidth for cloud connections, metro ethernet, and Data Centre backhaul. Offering extensive coverage in APAC and expanding globally, Megaport empowers users to manage their networks through its u...

Diversified Telecommunication Services
251-1K
Founded 2013
$26M raised

Description

  • Improve production reliability and system resilience within an SRE-scoped team.
  • Champion high standards of work and industry best practices.
  • Communicate with internal teams and stakeholders throughout requirements analysis, delivery, and demonstrations.
  • Investigate complex technical problems and contribute fresh ideas to improve outcomes.
  • Work across multiple technologies in a fast-changing environment.
  • Participate in on-call rotation, incident response, and blameless post-incident reviews.
  • Write code, handle alerts, improve solutions, and support teammates.
  • Collaborate across time zones in an asynchronous, globally distributed team.

Requirements

  • 5+ years administering Linux systems and related production infrastructure.
  • Collaborative SRE mindset with familiarity with SLIs, SLOs, SLAs, error budgets, blast radius, and blameless postmortems.
  • Strong focus on automation, toil reduction, and preventing problem recurrence.
  • Experience writing runbooks for a broader team.
  • Strong Kubernetes and ecosystem fundamentals.
  • Cloud infrastructure experience; AWS strongly preferred and bare-metal experience a bonus.
  • Strong scripting/tool development skills in Bash, plus either Python or Go preferred.
  • Infrastructure-as-code experience; Terraform preferred.
  • CI/CD and version control experience; GitHub preferred.
  • Database experience with Postgres, Cassandra, or ClickHouse preferred.
  • Experience operating production observability stacks across metrics, logs, and traces.
  • Comfort working on live production infrastructure with strong troubleshooting and incident-response ownership.
  • A history of continual professional development.
  • Self-directed and comfortable working with an async, globally distributed team and picking up adjacent work when needed.

Benefits

  • Remote-first flexible working environment with coworking options.
  • 4 weeks of paid annual leave, plus parental leave, birthday leave, and a purchased annual leave program.
  • Wellness allowance and employee wellbeing initiatives.
  • Generous study and training allowance plus 5 days of paid study leave.
  • Creative, modern workspaces for hybrid use.
  • Inclusive team environment with industry experts and fresh talent.
  • Recognition programs through Legend and Kudos awards.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Site Reliability Engineer

AssureSoft 51-250 Internet Software & Services

AssureSoft is hiring a remote Site Reliability Engineer to support production cloud infrastructure and platform reliability for long-term client projects.

Argo CD AWS Bash DNS Docker GCP GitHub Actions Go Grafana HTTP Kafka Kubernetes Linux Load Balancing Prometheus Python RabbitMQ Snowflake TCP/IP TLS TypeScript Unix
29 minutes ago

Site Reliability Engineer - Azure, Observability and Scripting

Jalasoft 1K-5K Internet Software & Services

Jalasoft is hiring a Site Reliability Engineer to support the reliability, scalability, and performance of cloud-native platforms on Microsoft Azure and Kubernetes.

Azure Bash Grafana Kubernetes OpenTelemetry Oracle PowerShell Prometheus Python Terraform
1 day ago

Senior Service Reliability Engineer

Thoughtworks 10K-50K Professional Services

Senior Service Reliability Engineer at Thoughtworks, focused on improving infrastructure reliability, observability, and incident response for customer-facing production systems.

Azure Bash Datadog ELK Stack GitOps Go Grafana Kubernetes Microservices Network Security New Relic Nomad Python REST API Serverless Terraform
1 day, 1 hour ago

Senior Service Reliability Engineer

Thoughtworks 10K-50K Professional Services

Thoughtworks is hiring a Senior Service Reliability Engineer to lead reliability-focused infrastructure work for production systems, improving resilience, observability, incident response, and operational efficiency in support of customer and business goals.

AWS Azure Bash CI/CD Datadog ELK Stack GCP GitOps Go Grafana Kubernetes Microservices New Relic Nomad Python REST API Serverless
1 day, 1 hour ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers