Resilient Co

Resilient Co

Resilient Co is a technology consulting company that empowers businesses with smart solutions and diverse teams, offering resilient support in the dynamic tech industry.

Professional Services
11-50
Founded 2020

Description

  • Design, deploy, and maintain production Kubernetes clusters and related services.
  • Build and maintain Python-based automation and operational tooling.
  • Integrate and operate Prometheus for monitoring, alerting, and observability.
  • Deploy and manage Ceph storage solutions for distributed workloads.
  • Support platform modernization initiatives and migrate services to cloud-native patterns.
  • Troubleshoot and resolve issues across compute, storage, and network layers in distributed systems.
  • Collaborate with development, SRE, and operations teams to define platform requirements and SLAs.
  • Document platform designs, runbooks, and operational procedures.
  • Participate in on-call rotations and incident response to maintain platform availability.

Requirements

  • 5+ years of experience in platform, infrastructure, or site reliability engineering roles.
  • Proven experience deploying and operating Kubernetes in production.
  • Strong Python skills for automation, tooling, and operational scripts.
  • Experience implementing and operating Prometheus-based monitoring and alerting.
  • Hands-on experience with Ceph or similar distributed storage systems.
  • Cloud experience with AWS and Azure, including designing, deploying, and operating services.
  • Demonstrated ability to troubleshoot distributed systems and resolve production incidents.
  • Experience collaborating across teams to deliver platform improvements and migrations.
  • Experience with OpenSearch (preferred).
  • Proficiency with Bash scripting (preferred).
  • Familiarity with Java-based services (preferred).
  • Experience with Fluent Bit for log collection (preferred).
  • Experience working with PostgreSQL (preferred).

Benefits

  • 12-month or longer engagement.
  • PST working hours (8:00 AM - 5:00 PM).
  • Client USA holiday calendar is mandatory.
  • No overtime required.
  • BYOD laptop policy.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Infrastructure Deployment Manager

Sphinx Defense 1-10 Construction & Engineering

Sphinx Defense is hiring an Infrastructure Deployment Manager to build and maintain production-ready infrastructure for software development environments supporting national security programs in Space.

Ansible Kubernetes SaltStack Terraform
1 hour, 46 minutes ago

Enterprise Platform Engineer V

Vatica Health 251-1K Internet Software & Services

Vatica Health is hiring an Enterprise Platform Engineer V to define and drive enterprise cloud architecture across a large multi-cloud infrastructure supporting the company’s growth.

AWS Azure CI/CD HIPAA Sentinel SIEM Terraform
2 hours, 16 minutes ago

Cloud Systems Engineer II

Take-Two Interactive Software 5K-10K Internet Software & Services

Take-Two Interactive is hiring a Cloud Systems Engineer to build and support scalable cloud and infrastructure systems across public cloud and on-prem environments for internal teams and customers.

Ansible AWS CI/CD Datadog DNS Docker EC2 GCP GitHub GitHub Actions Grafana Kubernetes Linux Load Balancing Prometheus Puppet Python TCP/IP Terraform
2 hours, 16 minutes ago

Infrastructure Consultant

Phoenix Software 251-1K IT Services

Phoenix is hiring an Infrastructure Consultant to deliver on-premise and hybrid infrastructure solutions for UK public sector customers across modernisation and support projects.

Active Directory DHCP DNS Network Security PowerShell Windows Server
2 hours, 31 minutes ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers