Capital.com

Capital.com

Capital.com is a leading fintech company providing online trading services through a smart investment app, offering access to 3700+ global markets with AI-powered features for secure and efficient trading.

Capital Markets
251-1K
Founded 2016
$25M raised

Description

  • Design, deploy, and maintain scalable cloud infrastructure on AWS with high availability, performance, and security.
  • Own and evolve Kubernetes cluster management, including bare-metal deployments, and support reliable containerised workloads with Docker and Helm.
  • Build and maintain CI/CD pipelines using GitLab CI and GitOps workflows with FluxCD or ArgoCD.
  • Define, manage, and review Infrastructure as Code using Terraform.
  • Lead monitoring and observability efforts, including dashboards, alerting, and log pipelines with VictoriaMetrics/Prometheus, Grafana, and the ELK stack.
  • Operate and optimize Apache Kafka ecosystems, including Strimzi, Kafka Connect, and MirrorMaker.
  • Drive incident response, root cause analysis, and post-mortem practices to improve reliability.
  • Collaborate with Engineering, Security, and Product teams to embed DevOps best practices across the organisation.
  • Mentor and guide junior engineers to raise the engineering bar for infrastructure reliability and automation.

Requirements

  • 6+ years of hands-on experience in a DevOps or SRE role.
  • Strong knowledge of AWS services, including VPC, EC2, EKS, S3, ECR, EBS, RDS, ElastiCache, IAM, KMS, Secrets Manager, SSM Parameter Store, CloudWatch, MSK, SNS, SQS, Route 53, Direct Connect, Transit Gateway, and ELB/ALB/NLB.
  • Solid Linux administration skills with a deep understanding of system internals.
  • Deep expertise in Kubernetes, including bare-metal cluster deployment and day-2 operations.
  • Proficiency with Docker and Helm.
  • Hands-on experience with Terraform as a primary Infrastructure as Code tool, including writing, reviewing, and maintaining production-grade modules.
  • Proven experience with GitLab CI for building and maintaining CI/CD pipelines; familiarity with GitOps practices using FluxCD or ArgoCD.
  • Strong background in monitoring and observability with VictoriaMetrics or Prometheus, Grafana, and the ELK stack.
  • Experience operating and managing Apache Kafka ecosystems, including Strimzi, Kafka Connect, and MirrorMaker.
  • Experience with Ansible for configuration management; AWX experience is a plus.
  • Proficiency in scripting and automation with Bash, Python, and Go.
  • Strong communication skills and the ability to collaborate cross-functionally in a fast-paced, regulated environment.
  • English language proficiency.

Benefits

  • Competitive salary.
  • Hybrid work arrangement with flexibility to work remotely.
  • Generous annual leave.
  • Employee referral program.
  • Comprehensive health and pension benefits, including medical insurance and pension plans.
  • 30 extra days per year to work remotely from anywhere in the world, subject to restrictions.
  • Two additional paid volunteer days each year.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Site Reliability Engineer 2 (Azure)

PhonePe 5K-10K Capital Markets

PhonePe Limited is hiring a Site Reliability Engineer to manage and scale core cloud infrastructure for a high-volume digital payments environment in India.

Ansible Azure Bash DNS Docker Go Grafana HAProxy InfluxDB Java Linux MySQL Nginx Prometheus Python RabbitMQ SaltStack Terraform Ubuntu
17 hours, 47 minutes ago

DevOps Engineer

Launchpad Technologies 51-250 Internet Software & Services

Launchpad Technologies Inc. is building a talent pool for remote DevOps opportunities focused on cloud infrastructure, CI/CD, and reliable, secure environments for clients across North America and beyond.

Agile Ansible AWS Azure Bash CloudFormation Datadog DevSecOps Docker GCP Git GitHub GitLab Grafana Kubernetes Microservices PowerShell Prometheus Python Serverless Terraform
18 hours, 2 minutes ago

Sr. Control System Engineer/Site Reliability Engineer (SRE)

QuEra Computing 11-50 Internet Software & Services

QuEra is seeking a Sr. Control System Engineer/Site Reliability Engineer to integrate and maintain the hardware and software systems that support its quantum control stack and keep development and production environments reliable.

Ansible Bash CI/CD Debian DHCP DNS Docker ELK Stack Embedded Systems Git GitLab CI Go Grafana Jenkins Kubernetes Linux Prometheus Python TCP/IP Terraform Ubuntu
18 hours, 17 minutes ago

Incident Commander

PENN Entertainment 10K-50K Hotels, Restaurants & Leisure

PENN Interactive is hiring an Incident Commander to join its site reliability team and lead cross-functional incident response for its online and physical platforms.

Ansible AWS Docker Elasticsearch GCP Helm JIRA Kafka Kubernetes Linux MySQL PostgreSQL Prometheus Python Redis Terraform
18 hours, 32 minutes ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers