Remote

Remote

Global HR Solutions & Employment Tools for Distributed Teams | Remote Hire international talent in minutes. Remote is the most disruptive global payroll, tax, HR and compliance solution for distributed teams. The easier way to employ internationally 🌍....

Professional Services
251-1K
Founded 2019
$496M raised

Description

  • Lead the discovery and delivery of reliability and infrastructure solutions for complex, ambiguous problems.
  • Own planning and execution of features and projects within the SRE/Platform domain.
  • Contribute to platform architecture, tooling, and roadmap decisions.
  • Define and operate reliability practices such as SLOs, SLIs, error budgets, alerting, and observability.
  • Resolve cross-team requests, identify systemic issues, and turn recurring issues into reusable fixes and runbooks.
  • Build and operationalize AI-native workflows, reusable prompts, skills, and tooling for the team.
  • Establish secure-by-default patterns, CI protections, and AI-assisted review practices.
  • Mentor less-senior engineers and provide timely, actionable feedback.
  • Participate in hiring, onboarding, and RFC discussions.
  • Collaborate with Security on platform hardening, threat mitigation, capacity, and cost-efficiency.
  • Participate in incident response and on-call rotations to maintain system reliability.

Requirements

  • Solid professional experience in SRE, DevOps, or Platform Engineering.
  • Hands-on experience operating and scaling Kubernetes production clusters and Docker/container tooling.
  • Experience building and managing cloud infrastructure on AWS or a similar cloud provider.
  • Strong infrastructure-as-code experience with Terraform.
  • Experience with reliability frameworks including SLOs, SLIs, error budgets, and alerting strategies.
  • Solid observability experience with OpenTelemetry, Grafana, Prometheus, or similar tools.
  • Experience with CI/CD and deployment automation, such as GitLab CI or GitHub Actions.
  • Comfort with Golang and Bash/scripting; broader programming experience is a plus.
  • Practical, embedded use of AI in infrastructure, operations, or development work with observable results.
  • Clear communication skills in an async-first, global environment.
  • Proactive, curious, and comfortable taking ownership of challenges.
  • Collaborative and respectful across cultures, time zones, and backgrounds.
  • Experience with one backend programming language such as Elixir, Node.js, or Python is preferred.
  • Experience running and configuring Linux systems in a non-cloud environment is preferred.
  • Security knowledge from both defensive and offensive perspectives is preferred.
  • Must submit application and CV in English.
  • Must upload a PDF CV or provide an up-to-date LinkedIn profile.

Benefits

  • Annual salary range of $53,300 to $119,850 USD.
  • Fair, unbiased compensation with equity pay.
  • Stock options.
  • Work from anywhere with a fully remote setup.
  • Flexible paid time off.
  • Flexible working hours in an async work environment.
  • 16 weeks of paid parental leave.
  • Mental health support services.
  • Learning budget.
  • Home office budget and IT equipment.
  • Budget for local in-person social events or co-working spaces.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

PENN Entertainment 10K-50K Hotels, Restaurants & Leisure

PENN Entertainment is seeking a Senior Site Reliability Engineer to build and operate the cloud infrastructure supporting large-scale sports betting and media platforms across regulated jurisdictions.

Argo CD AWS Bash CI/CD Datadog GCP GitHub Actions GitOps Go Helm Kubernetes Linux PostgreSQL Python Shell Scripting Terraform
16 hours, 54 minutes ago

Staff Site Reliability Engineer, Ads

Reddit 1K-5K Internet Software & Services

Reddit is hiring a Staff Site Reliability Engineer to provide technical leadership for reliability, scalability, and operational excellence across its advertising infrastructure and revenue-critical systems.

Apache Spark ClickHouse GCP Go Kafka Kubernetes
17 hours, 24 minutes ago

Sr Lead Network Reliability Engineer

Coupa Software 1K-5K Internet Software & Services

Coupa is hiring a Sr. Lead Network Development Engineer to scale and operate its global SaaS platform’s cloud networking infrastructure through automation, reliability engineering, and technical leadership.

Ansible AWS Azure Chef DNS Fortinet Go Java Kubernetes Linux Python Ruby TCP/IP Terraform TLS
1 day, 17 hours ago

Senior Monitoring/Observability Architect

Makpar 51-250 Internet Software & Services

Makpar is seeking a Senior Monitoring/Observability Architect to lead enterprise monitoring strategy, architecture, and implementation guidance for a large federal government program.

Datadog Splunk
1 day, 17 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers