Capital.com

Capital.com

Capital.com is a leading fintech company providing online trading services through a smart investment app, offering access to 3700+ global markets with AI-powered features for secure and efficient trading.

Capital Markets
251-1K
Founded 2016
$25M raised

Description

  • Own the full observability stack for metrics, logs, and traces, from pipeline design through day-2 operations.
  • Architect and operate the VictoriaMetrics cluster topology, including scraping, remote write, alerting rules, and cardinality control.
  • Operate OpenSearch clusters, including index lifecycle management, hot-warm-cold architecture, shard tuning, and ingest pipelines.
  • Build and maintain OpenTelemetry Collector pipelines and instrument services across Java, Python, and JavaScript/TypeScript stacks.
  • Run Kafka as the telemetry transport layer, including topic design, partition strategy, lag monitoring, and throughput tuning.
  • Manage log shipping infrastructure with Fluent Bit, Vector, or Fluentd and define structured logging standards across services.
  • Build Grafana dashboards and alerting that are clear, actionable, and useful for engineering teams.
  • Improve sampling, batching, and context propagation strategies across distributed services.
  • Participate in incident response, post-mortems, and reliability improvements driven by observability signals.
  • Mentor engineers on observability practices, tooling, and structured logging standards.

Requirements

  • 6+ years of experience in DevOps, SRE, or platform engineering roles.
  • At least 2 years of experience focused on observability tooling at production scale.
  • Deep hands-on experience with VictoriaMetrics or Prometheus, including MetricsQL/PromQL, exporters, service discovery, remote write, downsampling, and retention management.
  • Solid OpenSearch or Elasticsearch experience, including cluster operations, Query DSL, ISM policies, and ingest pipeline design.
  • Production experience with OpenTelemetry, including Collector configuration, OTLP, context propagation, and instrumentation across multiple languages.
  • Strong Kafka experience, including producer/consumer patterns, consumer group management, Kafka Connect, Schema Registry, and JMX-based monitoring.
  • Experience with Strimzi is a plus for running Kafka on Kubernetes.
  • Proficiency with log shippers such as Fluent Bit, Vector, or Fluentd and structured log parsing/normalization.
  • Working knowledge of Kubernetes, Helm, Argo CD/GitOps, Terraform, and Ansible.
  • Comfort in a hybrid AWS and on-prem environment, with solid networking knowledge as it applies to scraping and shipping pipelines.
  • Scripting ability in Bash or Python for automation and tooling.
  • Strong communication skills and English proficiency.

Benefits

  • Competitive salary.
  • Flexible work-life harmony with a hybrid work setup.
  • Generous annual leave.
  • Employee referral program.
  • Comprehensive health and pension benefits, including medical insurance.
  • 30 extra days to work remotely from anywhere in the world, with some restrictions.
  • Two additional paid volunteer days each year.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

PandaDoc 251-1K Internet Software & Services

PandaDoc is hiring a Site Reliability Engineer to help keep its document workflow platform reliable, resilient, and available for customers.

Agile AWS Django Grafana Java Kafka Kubernetes NATS PostgreSQL Python RabbitMQ Spring Boot
16 hours, 58 minutes ago

Senior Site Reliability Engineer

PandaDoc 251-1K Internet Software & Services

PandaDoc is hiring a Site Reliability Engineer to help keep its document workflow platform reliable through incident management, observability, production support, and resilience work across services.

Agile AWS Django Grafana Java Kafka Kubernetes NATS PostgreSQL Python RabbitMQ Spring Boot
16 hours, 59 minutes ago

Senior Site Reliability Engineer

PandaDoc 251-1K Internet Software & Services

PandaDoc is hiring a Site Reliability Engineer to keep its document workflow platform highly available and resilient while supporting production operations and reliability improvements.

Agile AWS Django Grafana Java Kafka Kubernetes NATS PostgreSQL Python RabbitMQ Spring Boot
1 day, 16 hours ago

Senior Site Reliability Engineer

PandaDoc 251-1K Internet Software & Services

PandaDoc is hiring a Site Reliability Engineer to help keep its production document workflow platform reliable, resilient, and low-downtime for customers.

AWS Django Grafana Java Kafka Kubernetes NATS PostgreSQL Python RabbitMQ Spring Boot
1 day, 16 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers