RemoteWoman

RemoteWoman

RemoteWoman is a platform that connects women with remote job opportunities at female-friendly companies in various fields such as marketing, development, design, product, sales, and support. They also offer a service for obtaining a legal US marriage ...

Internet Software & Services
1-10

Description

  • Lead reliability and observability strategy using Prometheus, Grafana, and ELK Stack.
  • Define actionable SLIs and SLOs aligned with business metrics.
  • Automate infrastructure and operational workflows with Terraform and Ansible across AWS, GCP, and Azure.
  • Optimize CI/CD pipelines for secure, fast, and zero-downtime releases.
  • Lead high-priority incident response, blameless post-mortems, and preventative improvements.
  • Partner with product, engineering, and platform teams to embed SRE practices into roadmaps.
  • Identify performance bottlenecks and evaluate technologies such as eBPF and container orchestration.
  • Participate in a four-week engineering and operations rotation, including hands-on troubleshooting and reliability work.
  • Participate in on-call coverage one week every 4–5 weeks, including weekend shifts from 02:00–10:00 UTC.

Requirements

  • 5+ years of experience in SRE, cloud operations, or DevOps, owning reliability for production platforms at scale.
  • Strong proficiency in Go or Python for automation tools, custom controllers, or SRE platform components.
  • Advanced knowledge of Linux internals, kernel parameters, networking, performance profiling, and troubleshooting.
  • Deep experience with AWS, GCP, Azure, or OpenStack and declarative infrastructure tools such as Terraform.
  • Experience building custom tooling with cloud SDKs and managing distributed systems.
  • Ability to anticipate operational risks, make architectural trade-offs, and lead infrastructure initiatives independently.
  • Strong cross-functional communication and collaboration skills.
  • Experience with custom orchestration, edge, storage, or operational tooling is preferred.
  • Experience managing Docker and production Kubernetes environments is preferred.
  • Familiarity with PaaS architectures or developer-facing cloud platforms is preferred; candidates must be legally authorized to work in Western Australia, and all roles require background checks.

Benefits

  • Remote-first work with a global, diverse team.
  • Flexible PTO and inclusive parental leave.
  • Company stock options.
  • Professional development, wellness, and office equipment budgets.
  • Internet reimbursement and a remote work travel program.
  • Annual team gatherings.
  • Inclusive, flexible workplace with accommodations available throughout the hiring process.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Arquitecto de soluciones / SRE - Argentina

Coderio 51-250 Internet Software & Services

Coderio busca un/a Arquitecto/a de Soluciones / SRE en Argentina para un entorno bancario, responsable de configurar, monitorear y analizar soluciones orientadas a pruebas de performance.

AWS Grafana Java Jenkins Kafka Kubernetes Microservices .NET Node.js OpenShift Prometheus Python RabbitMQ
1 day, 1 hour ago

SRE Engineer Contractor

Beauty For All Industries (BFA) 51-200 information technology & services

IPSY is seeking a remote Site Reliability Engineer Contractor in Mexico or Colombia, covering the PST time zone, to improve the availability, resilience, and operational reliability of its beauty membership platform.

Amplitude AWS Bash CI/CD Contentful Datadog Grafana Microservices Netlify New Relic OpsGenie PagerDuty Prometheus Python Terraform
3 days, 1 hour ago

[Job 31894] AI ORCHESTRATOR (APP SRE)

CI&T 5K-10K Internet Software & Services

A CI&T busca um AI Orchestrator para governar a entrega ponta a ponta de um projeto de monitoramento SRE em jornadas de aplicativos de saúde, coordenando equipes e agentes de IA desde o upstream até a produção.

Agile
5 days, 2 hours ago

Senior Site Reliability Engineer (SRE)

UJET 251-1K Professional Services

UJET is seeking a Senior Site Reliability Engineer to build and scale its SRE function, improving reliability, reducing operational toil, and establishing production engineering best practices for its cloud-based contact center platform.

AWS Azure Go Java Kubernetes Microservices Python Terraform
6 days, 1 hour ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers