Resil

Resil

Resil specializes in providing AI-powered supply chain risk management solutions that enable organizations to detect threats in real time and take proactive measures to enhance supply chain resiliency.

Internet Software & Services
251-1K
Founded 2010

Description

  • Design, implement, and manage scalable, highly available systems on Azure Cloud.
  • Monitor system performance, troubleshoot issues, and ensure uptime and reliability.
  • Manage and optimize Kubernetes clusters and containerized workloads using Docker.
  • Build and maintain CI/CD pipelines using GitHub Actions and related tools.
  • Implement infrastructure as code and deployment automation using Helm Charts.
  • Work with distributed systems including Kafka, Redis, PostgreSQL, and Hadoop/HDFS.
  • Configure and manage Cloudflare for performance, security, and traffic routing.
  • Set up monitoring, alerting, and observability using tools such as Grafana.
  • Collaborate with development teams to improve reliability and deployment practices.
  • Perform root cause analysis and implement preventive measures.
  • Ensure security best practices and compliance across infrastructure.

Requirements

  • 6–12 years of experience in SRE, DevOps, or related roles.
  • Strong hands-on experience with Azure Cloud services.
  • Solid experience in Linux system administration.
  • Expertise in Docker and Kubernetes, including deployment, scaling, and troubleshooting.
  • Experience with Kafka, Redis, and PostgreSQL.
  • Working knowledge of the Hadoop ecosystem, including HDFS.
  • Experience with Cloudflare for CDN, security, and DNS management.
  • Proficiency with CI/CD tools, especially GitHub and GitHub Actions.
  • Experience with Helm Charts and Kubernetes deployments.
  • Strong understanding of monitoring and logging tools such as Grafana.
  • Experience with large-scale distributed systems is preferred.
  • Knowledge of infrastructure automation tools such as Terraform or Ansible is preferred.
  • Exposure to security and compliance best practices is preferred.
  • Experience with Databricks, ClickHouse, or MLOps is preferred.
  • Strong problem-solving and troubleshooting skills.

Benefits

  • Fully remote work environment with opportunities to connect in person.
  • A collaborative, innovation-driven culture with ownership and purpose.
  • Opportunities for technical growth and a voice in shaping impactful technology.
  • Exposure to cutting-edge cloud and distributed systems.
  • Exposure to large-scale, high-impact platforms.
  • Full-stack benefits covering health, wealth, and wellbeing.
  • Equal opportunity employer status.
  • Support for applicants with disabilities during the application process.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Technical Support Engineer (GPU Clusters) - US Weekends

Together 1-10 IT Services

Together AI is hiring a Technical Support Engineer to support customers building and operating AI training, fine-tuning, and inference systems on Kubernetes GPU infrastructure.

Ansible Kubernetes Machine Learning
14 hours ago

Staff Site Reliability Engineer

BeyondTrust 1K-5K Professional Services

BeyondTrust is hiring a Staff Site Reliability Engineer to lead the evolution of its Password Safe platform, infrastructure, and deployment ecosystem across cloud and on-premises environments.

Ansible AWS Azure C# CI/CD Datadog DevSecOps Docker GitOps Go Java Kubernetes Linux Microservices OpenTelemetry Secrets Management Terraform
14 hours ago

Senior Site Reliability Engineer

Alpaca 51-250 Capital Markets

Alpaca is hiring a Site Reliability Engineer to keep its brokerage platform reliable, observable, and operable across cloud infrastructure, Kubernetes, and PostgreSQL on the trading-critical path.

DNS GitOps Go Kafka Kubernetes Linux Load Balancing PostgreSQL Python RabbitMQ Secrets Management TLS
1 day, 13 hours ago

Software/Site Reliability Engineer - FedRAMP

Tenable 1K-5K Internet Software & Services

Tenable is hiring a Site Reliability Engineer to help scale and operate its cloud-based vulnerability management platform for private and U.S. Government cloud customers.

Agile AWS Azure Bash CI/CD Datadog Docker DynamoDB Elasticsearch GCP Go Gradle Groovy Helm Java Kafka Kotlin Kubernetes Microservices Node.js OpenSearch OpenTelemetry Python Splunk Terraform
3 days, 13 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers