Remote

Remote

Global HR Solutions & Employment Tools for Distributed Teams | Remote Hire international talent in minutes. Remote is the most disruptive global payroll, tax, HR and compliance solution for distributed teams. The easier way to employ internationally 🌍....

Professional Services
251-1K
Founded 2019
$496M raised

Description

  • Design, implement, and maintain infrastructure-as-code patterns using Terraform and Kubernetes.
  • Build and maintain monitoring, logging, and alerting systems for production services.
  • Lead incident response, conduct post-mortems, and drive reliability improvements.
  • Work with the Security team to embed security and compliance into Build infrastructure.
  • Continuously optimize system performance, resource utilization, and cloud costs.
  • Eliminate manual operational toil through automation, tools, and improved processes.
  • Partner with platform teams to improve API, MCP, and CLI resilience and observability.
  • Provide infrastructure feedback that helps shape platform evolution and developer experience.

Requirements

  • Senior-level experience in Site Reliability Engineering, DevOps Engineering, or SysOps roles.
  • Experience standing up and operating production systems at scale.
  • Deep hands-on experience running Kubernetes in production.
  • Solid AWS fundamentals across compute, networking, storage, and managed services.
  • Proficiency with Terraform or similar infrastructure-as-code tools.
  • Experience with CI/CD and deployment automation tools such as GitLab, GitHub Actions, or Jenkins.
  • Strong bash scripting and comfort debugging system-level issues and logs.
  • Understanding of Linux kernel basics.
  • Ability to communicate complex infrastructure decisions clearly to technical and non-technical stakeholders.
  • Experience with at least one backend programming language such as Elixir, Python, Go, Java, or Node.js is a plus.
  • Experience in consultancy settings is a plus.
  • Experience with container registries and artifact management such as ECR or Docker Hub is a plus.
  • Experience with observability tools such as Datadog, Prometheus, ELK, or Grafana is a plus.
  • Experience building or scaling multi-tenant platforms is a plus.
  • Application materials must be submitted in English, and a PDF CV or LinkedIn profile is required.

Benefits

  • Annual salary range of $54,000 to $150,000 USD.
  • Fair, unbiased compensation with equity pay.
  • Stock options.
  • Work from anywhere.
  • Flexible paid time off.
  • Flexible working hours in an async environment.
  • 16 weeks of paid parental leave.
  • Mental health support services.
  • Learning budget.
  • Home office budget and IT equipment.
  • Budget for local in-person social events or co-working spaces.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Application Reliability Engineer

Innodata 1K-5K IT Services

Innodata is hiring a hands-on Application Support Engineer to maintain and enhance business-critical enterprise applications running on Google App Engine and microservices across production and release environments.

API Gateway Apigee CI/CD Docker Firestore GCP Generative AI GitHub Actions GitLab CI Go gRPC Java JavaScript Jenkins Kubernetes LLM Load Balancing Looker Microservices Node.js Python REST API SQL Tableau Terraform Vertex AI
6 hours, 17 minutes ago

Sr. Reliability Engineer

Syner-G BioPharma 51-250 Professional Services

Syner-G is hiring a Senior Reliability Engineer to lead reliability and maintenance optimization for critical building and facility systems supporting life sciences and research operations.

6 hours, 32 minutes ago

SRE and Devops Team Manager

Eltropy 51-250 Communications Equipment

Eltropy is hiring a remote SRE and DevOps Team Engineer in India to lead reliability, automation, and cloud operations for its SaaS platform.

AWS GCP GitHub Actions Jenkins Kubernetes Pulumi Terraform
1 day, 6 hours ago

Lead Site Reliability Engineer - Active Directory

Coupa Software 1K-5K Internet Software & Services

Coupa is seeking a Lead Site Reliability Engineer to define and operate its global Active Directory and identity infrastructure within the Cloud Operations team.

Active Directory Ansible AWS Azure Chef CI/CD DHCP DNS GCP GitOps New Relic PowerShell Python SIEM TCP/IP Terraform
1 day, 6 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers