Resilient Co

Resilient Co

Resilient Co is a technology consulting company that empowers businesses with smart solutions and diverse teams, offering resilient support in the dynamic tech industry.

Professional Services
11-50
Founded 2020

Description

  • Design, deploy, and maintain production Kubernetes clusters and related services.
  • Build and maintain Python-based automation and operational tooling.
  • Integrate and operate Prometheus for monitoring, alerting, and observability.
  • Deploy and manage Ceph storage solutions for distributed workloads.
  • Support platform modernization initiatives and migrate services to cloud-native patterns.
  • Troubleshoot and resolve issues across compute, storage, and network layers in distributed systems.
  • Collaborate with development, SRE, and operations teams to define platform requirements and SLAs.
  • Document platform designs, runbooks, and operational procedures.
  • Participate in on-call rotations and incident response to maintain platform availability.

Requirements

  • 5+ years of experience in platform, infrastructure, or site reliability engineering roles.
  • Proven experience deploying and operating Kubernetes in production.
  • Strong Python skills for automation, tooling, and operational scripts.
  • Experience implementing and operating Prometheus-based monitoring and alerting.
  • Hands-on experience with Ceph or similar distributed storage systems.
  • Cloud experience with AWS and Azure, including designing, deploying, and operating services.
  • Demonstrated ability to troubleshoot distributed systems and resolve production incidents.
  • Experience collaborating across teams to deliver platform improvements and migrations.
  • Experience with OpenSearch (preferred).
  • Proficiency with Bash scripting (preferred).
  • Familiarity with Java-based services (preferred).
  • Experience with Fluent Bit for log collection (preferred).
  • Experience working with PostgreSQL (preferred).

Benefits

  • 12-month or longer engagement.
  • PST working hours (8:00 AM - 5:00 PM).
  • Client USA holiday calendar is mandatory.
  • No overtime required.
  • BYOD laptop policy.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Infrastructure Engineer

Sysdig 251-1K IT Services

As a Site Reliability/Infrastructure Engineer at Sysdig, you will build and operate multi-cloud and on-premise infrastructure while improving the reliability, scalability, security, and performance of production systems.

AWS Azure Bash Docker Go Kubernetes Linux Microservices Python
10 minutes ago

Cloud Security Engineer

Smile Digital Health 251-1K IT Services

Smile Digital Health is seeking a Cloud Security Engineer to design, automate, deploy, and support secure production-grade infrastructure for healthcare data platforms across AWS, Azure, OCI, and GCP.

Ansible AWS Azure Docker HIPAA Kubernetes OpenShift Terraform
1 day, 23 hours ago

Director of Infrastructure

Sysdig 251-1K IT Services

Sysdig is hiring an infrastructure leader to own global cloud strategy, reliability, cost efficiency, and team execution for its large-scale, multi-tenant cloud security platform.

AWS Azure Kubernetes Pulumi Terraform
1 day, 23 hours ago

Cloud Infrastructure Engineer (Remote LATAM)

Atmosera 51-250 IT Services

Atmosera is seeking a senior Azure Cloud Infrastructure Engineer contractor to design, implement, migrate, and operate secure cloud environments while advising clients and supporting internal operations.

Agile Ansible Azure PowerShell SQL Server Terraform Windows Server
3 days, 23 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers