Resilient Co

Resilient Co

Resilient Co is a technology consulting company that empowers businesses with smart solutions and diverse teams, offering resilient support in the dynamic tech industry.

Professional Services
11-50
Founded 2020

Description

  • Design, deploy, and maintain production Kubernetes clusters and related services.
  • Build and maintain Python-based automation and operational tooling.
  • Integrate and operate Prometheus for monitoring, alerting, and observability.
  • Deploy and manage Ceph storage solutions for distributed workloads.
  • Support platform modernization initiatives and migrate services to cloud-native patterns.
  • Troubleshoot and resolve issues across compute, storage, and network layers in distributed systems.
  • Collaborate with development, SRE, and operations teams to define platform requirements and SLAs.
  • Document platform designs, runbooks, and operational procedures.
  • Participate in on-call rotations and incident response to maintain platform availability.

Requirements

  • 5+ years of experience in platform, infrastructure, or site reliability engineering roles.
  • Proven experience deploying and operating Kubernetes in production.
  • Strong Python skills for automation, tooling, and operational scripts.
  • Experience implementing and operating Prometheus-based monitoring and alerting.
  • Hands-on experience with Ceph or similar distributed storage systems.
  • Cloud experience with AWS and Azure, including designing, deploying, and operating services.
  • Demonstrated ability to troubleshoot distributed systems and resolve production incidents.
  • Experience collaborating across teams to deliver platform improvements and migrations.
  • Experience with OpenSearch (preferred).
  • Proficiency with Bash scripting (preferred).
  • Familiarity with Java-based services (preferred).
  • Experience with Fluent Bit for log collection (preferred).
  • Experience working with PostgreSQL (preferred).

Benefits

  • 12-month or longer engagement.
  • PST working hours (8:00 AM - 5:00 PM).
  • Client USA holiday calendar is mandatory.
  • No overtime required.
  • BYOD laptop policy.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Pillar Lead: Application & Endpoint Specialist (R-00211)

True Zero Technologies 11-50 Internet Software & Services

True Zero Technologies is seeking an Application & Infrastructure Specialist to support Zero Trust architecture and implementation across applications, workloads, cloud services, platforms, and legacy infrastructure while enabling practical enterprise modernization.

AWS Azure Kubernetes Secrets Management
1 day, 17 hours ago

[Job- 31345]Senior DevOps Engineer (Cloud Network Foundation), Brazil

CI&T 5K-10K Internet Software & Services

As an AWS Network Foundation Engineer at CI&T, you will co-own multi-account AWS landing-zone networking for enterprise clients, ensuring secure, tested connectivity, routing, inspection, and operational handover.

AWS Network Security Terraform
1 day, 17 hours ago

IT Infrastructure Administrator

Omilia 251-1K IT Services

Omilia is seeking an IT Infrastructure Administrator to operate and improve its on-premises, cloud, and hybrid infrastructure supporting internal operations and platform delivery.

Ansible AWS Azure Cisco Datadog DHCP DNS Docker Fortinet GCP Grafana Kubernetes Linux Terraform Windows Server Zabbix
1 day, 17 hours ago

IT Engineer, Internal AI Infrastructure

Figma 1K-5K Internet Software & Services

Figma is hiring an infrastructure-focused software engineer for its Internal AI team to build the shared hosting, routing, security, observability, and cost-management infrastructure that brings employee-facing AI applications from prototype to production.

System Design
2 days, 16 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers