Resil

Resil

Resil specializes in providing AI-powered supply chain risk management solutions that enable organizations to detect threats in real time and take proactive measures to enhance supply chain resiliency.

Internet Software & Services
251-1K
Founded 2010

Description

  • Design, implement, and manage scalable, highly available systems on Azure Cloud.
  • Monitor system performance, troubleshoot issues, and ensure uptime and reliability.
  • Manage and optimize Kubernetes clusters and containerized workloads using Docker.
  • Build and maintain CI/CD pipelines using GitHub Actions and related tools.
  • Implement infrastructure as code and deployment automation using Helm Charts.
  • Work with distributed systems including Kafka, Redis, PostgreSQL, and Hadoop/HDFS.
  • Configure and manage Cloudflare for performance, security, and traffic routing.
  • Set up monitoring, alerting, and observability using tools such as Grafana.
  • Collaborate with development teams to improve reliability and deployment practices.
  • Perform root cause analysis and implement preventive measures.
  • Ensure security best practices and compliance across infrastructure.

Requirements

  • 6–12 years of experience in SRE, DevOps, or related roles.
  • Strong hands-on experience with Azure Cloud services.
  • Solid experience in Linux system administration.
  • Expertise in Docker and Kubernetes, including deployment, scaling, and troubleshooting.
  • Experience with Kafka, Redis, and PostgreSQL.
  • Working knowledge of the Hadoop ecosystem, including HDFS.
  • Experience with Cloudflare for CDN, security, and DNS management.
  • Proficiency with CI/CD tools, especially GitHub and GitHub Actions.
  • Experience with Helm Charts and Kubernetes deployments.
  • Strong understanding of monitoring and logging tools such as Grafana.
  • Experience with large-scale distributed systems is preferred.
  • Knowledge of infrastructure automation tools such as Terraform or Ansible is preferred.
  • Exposure to security and compliance best practices is preferred.
  • Experience with Databricks, ClickHouse, or MLOps is preferred.
  • Strong problem-solving and troubleshooting skills.

Benefits

  • Fully remote work environment with opportunities to connect in person.
  • A collaborative, innovation-driven culture with ownership and purpose.
  • Opportunities for technical growth and a voice in shaping impactful technology.
  • Exposure to cutting-edge cloud and distributed systems.
  • Exposure to large-scale, high-impact platforms.
  • Full-stack benefits covering health, wealth, and wellbeing.
  • Equal opportunity employer status.
  • Support for applicants with disabilities during the application process.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

PENN Entertainment 10K-50K Hotels, Restaurants & Leisure

PENN Entertainment is seeking a Senior Site Reliability Engineer to build and operate the cloud infrastructure supporting large-scale sports betting and media platforms across regulated jurisdictions.

Argo CD AWS Bash CI/CD Datadog GCP GitHub Actions GitOps Go Helm Kubernetes Linux PostgreSQL Python Shell Scripting Terraform
16 hours, 41 minutes ago

Staff Site Reliability Engineer, Ads

Reddit 1K-5K Internet Software & Services

Reddit is hiring a Staff Site Reliability Engineer to provide technical leadership for reliability, scalability, and operational excellence across its advertising infrastructure and revenue-critical systems.

Apache Spark ClickHouse GCP Go Kafka Kubernetes
17 hours, 11 minutes ago

Sr Lead Network Reliability Engineer

Coupa Software 1K-5K Internet Software & Services

Coupa is hiring a Sr. Lead Network Development Engineer to scale and operate its global SaaS platform’s cloud networking infrastructure through automation, reliability engineering, and technical leadership.

Ansible AWS Azure Chef DNS Fortinet Go Java Kubernetes Linux Python Ruby TCP/IP Terraform TLS
1 day, 16 hours ago

Senior Monitoring/Observability Architect

Makpar 51-250 Internet Software & Services

Makpar is seeking a Senior Monitoring/Observability Architect to lead enterprise monitoring strategy, architecture, and implementation guidance for a large federal government program.

Datadog Splunk
1 day, 16 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers