CXM Direct

CXM Direct

CXM Direct: A reliable broker with advanced trading tools and innovative solutions for traders worldwide.

Capital Markets
51-250
Founded 2015

Description

  • Participate in the on-call rotation for production trading systems and lead incident response during service disruptions.
  • Investigate production incidents, perform root cause analysis, and implement preventive actions to reduce recurrence.
  • Build and maintain Grafana dashboards, Prometheus alerts, and operational health views across applications, infrastructure, and databases.
  • Instrument .NET services to improve telemetry, metrics, logging, and visibility into service health and customer impact.
  • Define, implement, and monitor SLIs, SLOs, and error budgets.
  • Troubleshoot issues across .NET/C# applications, Windows Server, Aurora PostgreSQL, AWS infrastructure, and CI/CD pipelines.
  • Improve deployment safety, release automation, and rollback strategies.
  • Partner with developers to improve application operability, resilience, and fault isolation.
  • Automate operational tasks through scripting and infrastructure automation.
  • Create and maintain runbooks, operational documentation, and incident response procedures.

Requirements

  • 3–5 years of experience in a mid-level role.
  • Strong experience debugging and supporting .NET/C# applications in production.
  • Hands-on experience with Windows Server environments.
  • Strong PowerShell scripting skills.
  • Experience with Python or Bash.
  • Experience with Grafana, Prometheus, and Loki or equivalent observability tools.
  • Strong understanding of metrics, logging, tracing, and alerting best practices.
  • Experience with modern CI/CD pipelines and deployment strategies.
  • Experience working with AWS.
  • Hands-on experience with Terraform or other Infrastructure as Code tools.
  • Experience troubleshooting and supporting Aurora PostgreSQL or other relational databases.
  • Practical experience with SLIs, SLOs, error budgets, incident response, RCA, alert design, and production operations.
  • Preferred: experience supporting high-availability or low-latency financial or trading systems.
  • Preferred: familiarity with MetaTrader environments or financial technology platforms.
  • Preferred: experience with distributed systems and microservices.
  • Preferred: knowledge of OpenTelemetry or similar observability frameworks.
  • Preferred: exposure to Docker, Kubernetes, or containerized environments.

Benefits

  • Remote work from the Americas, with LatAm preferred.
  • Working hours aligned to Americas time zones (UTC-3 to UTC-8).
  • Full-time, permanent employment.
  • On-call rotation aligned with the London trading day.
  • Opportunity to work on mission-critical trading infrastructure.
  • Chance to solve challenging reliability and scalability problems in a real-time environment.
  • Opportunity to influence reliability strategy and engineering best practices across the platform.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Site Reliability Engineer

VantageScore 11-50 Banks

Site Reliability Engineer at a growing engineering team, focused on DevSecOps for maintaining the reliability, security, and compliance of cloud infrastructure, APIs, and software supply chains.

Agile AWS AWS CDK Bash CI/CD CloudFormation CodePipeline Datadog DevSecOps Docker EC2 GitHub Actions Grafana HashiCorp Vault Kong Kubernetes Microservices Python REST API Scrum Terraform
5 hours, 3 minutes ago

Customer Reliability Engineer

iPiD 11-50 Internet Software & Services

iPiD is hiring a Customer Reliability Engineer to own production reliability, customer deployments, and operational excellence for its global KYP verification platform.

Ansible CI/CD GitOps Helm Kubernetes Linux Microservices Terraform
5 hours, 33 minutes ago

Site Reliability Engineer

CSC Generation 251-1K Internet Software & Services

Backcountry is hiring a Site Reliability Engineer in Costa Rica to keep its ecommerce platform reliable, scalable, and observable across a multi-cloud environment.

Ansible Argo CD AWS AWS CDK Bash CI/CD Docker GCP GitOps Grafana Helm Kubernetes Linux Node.js OpenSearch Prometheus Python Terraform TypeScript
1 day, 4 hours ago

Ssr Monitoring and Observability Analyst

Coderio 51-250 Internet Software & Services

Coderio is hiring an Observability & Monitoring Analyst to design and operate monitoring systems that improve availability, performance, and incident response across global clients’ IT environments.

AWS Azure Bash Datadog DNS Docker ELK Stack Fluentd GCP Grafana Jaeger Kibana Kubernetes Linux Load Balancing Logstash New Relic OpenTelemetry Prometheus Python TCP/IP Zipkin
1 day, 5 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers