Remote

Remote

Global HR Solutions & Employment Tools for Distributed Teams | Remote Hire international talent in minutes. Remote is the most disruptive global payroll, tax, HR and compliance solution for distributed teams. The easier way to employ internationally 🌍....

Professional Services
251-1K
Founded 2019
$496M raised

Description

  • Design, implement, and maintain infrastructure-as-code patterns using Terraform and Kubernetes for standard connectors and custom builds.
  • Build and maintain monitoring, logging, and alerting systems to support observability.
  • Lead incident response efforts, conduct post-mortems, and drive reliability improvements.
  • Work with the Security team to embed security into the Build infrastructure and support compliance across 100+ jurisdictions.
  • Continuously optimize system performance, resource utilization, and cloud costs.
  • Identify and eliminate manual operational toil through automation and improved processes.
  • Partner with platform teams to ensure APIs, MCP, and CLI are resilient and observable.
  • Provide infrastructure feedback that helps shape platform evolution and developer experience.

Requirements

  • Senior-level experience in Site Reliability Engineering, DevOps Engineering, or SysOps roles.
  • Experience standing up and operating production systems at scale.
  • Deep hands-on experience running Kubernetes in production.
  • Solid AWS fundamentals across compute, networking, storage, and managed services.
  • Proficiency with Terraform or similar infrastructure-as-code tools.
  • Experience with CI/CD and deployment automation tools such as GitLab, GitHub Actions, or Jenkins.
  • Strong bash scripting skills.
  • Comfort debugging system-level issues, reading logs, and understanding Linux kernel basics.
  • Ability to communicate complex infrastructure decisions clearly to technical and non-technical stakeholders.
  • Experience with at least one backend programming language such as Elixir, Python, Go, Java, or Node.js (preferred).
  • Experience in consultancy settings (preferred).
  • Experience with container registries and artifact management such as ECR or Docker Hub (preferred).
  • Experience with observability tools such as Datadog, Prometheus, ELK, or Grafana (preferred).
  • Experience working with or scaling multi-tenant platforms (preferred).

Benefits

  • Annual salary range of $54,000 to $150,000 USD.
  • Work from anywhere with fully remote employment.
  • Flexible paid time off.
  • Flexible working hours with an async work culture.
  • 16 weeks of paid parental leave.
  • Mental health support services.
  • Stock options.
  • Learning budget.
  • Home office budget and IT equipment.
  • Budget for local in-person social events or co-working spaces.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Site Reliability Engineer III

onXmaps 251-1K Food Products

onX is hiring a Site Reliability Engineer to manage the infrastructure, deployment automation, and observability that help developers ship reliably at scale for its outdoor technology products.

Apache Airflow CockroachDB GCP Kubernetes OpenTelemetry Prometheus SQL Terraform
4 hours, 36 minutes ago

Site Reliability Engineer - South Korea

MinIO 51-250 Internet Software & Services

MinIO is hiring a Site Reliability Engineer to help enhance and operate its cloud-native storage platform for high-performance, scalable, and durable data storage and retrieval.

C C++ GitOps Go Kubernetes Microservices Rust
4 hours, 36 minutes ago

Senior Site Reliability Engineer- FedRamp

Veeam Software 1K-5K Internet Software & Services

Veeam is hiring a Site Reliability Engineer to help build its global SRE function for the Veeam Data Cloud, focused on the Government and Sovereign Cloud environment.

Argo CD Azure Bitbucket C# CI/CD ELK Stack Git GitHub Actions GitLab CI GitOps Go Grafana HIPAA Java JavaScript Kubernetes OpenTelemetry Prometheus Pulumi Terraform TypeScript
5 hours, 36 minutes ago

Senior Site Reliability Engineer

RUNWARE 1-10 Internet Software & Services

Runware is hiring a Site Reliability Engineer to keep its AI inference and serverless GPU platform reliable, performant, and resilient as the business scales.

CDN ClickHouse Go Kubernetes Load Balancing Machine Learning MySQL PHP Python RabbitMQ Redis Serverless
6 hours, 6 minutes ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers