Remote

Remote

Global HR Solutions & Employment Tools for Distributed Teams | Remote Hire international talent in minutes. Remote is the most disruptive global payroll, tax, HR and compliance solution for distributed teams. The easier way to employ internationally 🌍....

Professional Services
251-1K
Founded 2019
$496M raised

Description

  • Design, implement, and maintain infrastructure-as-code patterns using Terraform and Kubernetes.
  • Build and maintain monitoring, logging, and alerting systems for production services.
  • Lead incident response, conduct post-mortems, and drive reliability improvements.
  • Work with the Security team to embed security and compliance into Build infrastructure.
  • Continuously optimize system performance, resource utilization, and cloud costs.
  • Eliminate manual operational toil through automation, tools, and improved processes.
  • Partner with platform teams to improve API, MCP, and CLI resilience and observability.
  • Provide infrastructure feedback that helps shape platform evolution and developer experience.

Requirements

  • Senior-level experience in Site Reliability Engineering, DevOps Engineering, or SysOps roles.
  • Experience standing up and operating production systems at scale.
  • Deep hands-on experience running Kubernetes in production.
  • Solid AWS fundamentals across compute, networking, storage, and managed services.
  • Proficiency with Terraform or similar infrastructure-as-code tools.
  • Experience with CI/CD and deployment automation tools such as GitLab, GitHub Actions, or Jenkins.
  • Strong bash scripting and comfort debugging system-level issues and logs.
  • Understanding of Linux kernel basics.
  • Ability to communicate complex infrastructure decisions clearly to technical and non-technical stakeholders.
  • Experience with at least one backend programming language such as Elixir, Python, Go, Java, or Node.js is a plus.
  • Experience in consultancy settings is a plus.
  • Experience with container registries and artifact management such as ECR or Docker Hub is a plus.
  • Experience with observability tools such as Datadog, Prometheus, ELK, or Grafana is a plus.
  • Experience building or scaling multi-tenant platforms is a plus.
  • Application materials must be submitted in English, and a PDF CV or LinkedIn profile is required.

Benefits

  • Annual salary range of $54,000 to $150,000 USD.
  • Fair, unbiased compensation with equity pay.
  • Stock options.
  • Work from anywhere.
  • Flexible paid time off.
  • Flexible working hours in an async environment.
  • 16 weeks of paid parental leave.
  • Mental health support services.
  • Learning budget.
  • Home office budget and IT equipment.
  • Budget for local in-person social events or co-working spaces.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

SITE RELIABILITY ENGINEER III

Harford County Public Library 51-250 Diversified Consumer Services

Site Reliability Engineer na Stone, atuando no time de Foundation Platform para fortalecer a plataforma interna de tecnologia com foco em observabilidade, automação e estabilidade dos sistemas.

Ansible Argo CD AWS Azure Datadog Docker GCP GitHub Actions Go Grafana Kubernetes Linux Node.js OpenTelemetry Prometheus Python Splunk Terraform
23 minutes ago

Sr. Site Reliability Engineer (Starlink)

SpaceX 10K-50K Aerospace & Defense

SpaceX is hiring a Sr. Site Reliability Engineer for Starlink to improve the reliability, scalability, and performance of the systems supporting its satellite internet service.

Apache Spark C# CI/CD Flink Git Go HDFS Java Kafka Kubernetes Linux Python Scala
38 minutes ago

Head of Platform Engineering

dLocal 251-1K Diversified Financial Services

dLocal is seeking a senior leader to own its engineering platform, reliability posture, and AI-assisted development transformation across a global payments business serving emerging markets.

CI/CD Microservices
53 minutes ago

Database Reliability Engineer

Alex Staff Agency 11-50 Professional Services

Senior Database Reliability Engineer for an infrastructure DBA team, responsible for keeping production database services reliable and automating operational work across a multi-database environment.

Ansible ClickHouse DNS Grafana Linux MongoDB OpsGenie PostgreSQL Redis Terraform TLS
1 hour, 23 minutes ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers