DRW

DRW

DRW is a diversified trading firm that leverages technology and risk management to identify and capture global trading and investment opportunities across a wide range of asset classes and markets.

Capital Markets
1K-5K
Founded 1992

Description

  • Deploy, maintain, and optimize GPU infrastructure for large-scale LLM inference workloads.
  • Provision, configure, and deploy GPU server fleets.
  • Architect and implement distributed serving solutions for multi-node, multi-GPU model deployments.
  • Manage GPU-enabled Kubernetes clusters for LLM and ML workloads.
  • Configure network infrastructure, including load balancers, firewalls, and inter-node communication for GPU clusters.
  • Implement and optimize storage solutions for model weights and inference caches.
  • Troubleshoot performance bottlenecks across hardware, drivers, networking, and the application layer.
  • Research and evaluate emerging GPU technologies, model serving frameworks, and infrastructure optimizations.
  • Collaborate with ML engineers to profile model performance and implement inference acceleration techniques.
  • Drive reliability improvements through monitoring, alerting, capacity planning, and incident response.

Requirements

  • Bachelor's or Master's degree in Computer Science, Systems Engineering, or a related field.
  • 5+ years of experience in DevOps, SRE, or infrastructure engineering roles.
  • Strong experience with GPU infrastructure, model serving frameworks such as vLLM or SGLang, and GPU driver management.
  • Hands-on experience optimizing deep learning workloads on GPU clusters for inference or training.
  • Deep Linux systems knowledge, including network configuration, storage optimization, and Kubernetes orchestration.
  • Experience with infrastructure-as-code tools such as Ansible or Terraform.
  • Strong understanding of distributed systems, networking protocols such as TCP/IP and HTTP/2, and load balancing.
  • Proficiency in Python and Bash scripting for automation.
  • Experience with monitoring and observability tools such as Prometheus and Grafana.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Behavior Planning Software Engineer - ADAS

Lucid Motors 1K-5K Automotive

Lucid Motors is hiring a Behavior Planning Engineer for its ADAS Autonomous Driving team to develop and bring planning and prediction software and algorithms for autonomous driving features into production.

C++ Machine Learning Python Reinforcement Learning
3 hours, 51 minutes ago

Senior Intelligent Automation Engineer

Roboyo 251-1K Electrical Equipment

Roboyo AI is hiring a Senior Automation Engineer to lead client delivery workstreams and build production-grade intelligent automation solutions within its consulting and technology services practice.

Java Python
4 hours, 6 minutes ago

Detections Engineer

Shift5 51-250 Airlines

Shift5 is hiring a Detection Engineer to develop proof-of-concept detection capabilities for onboard operational technology across cybersecurity, predictive maintenance, and telemetry anomaly use cases.

C C++ Cybersecurity Docker Embedded Systems Git GitHub Actions Linux Machine Learning Python Rust
4 hours, 6 minutes ago

Manager / Sr. Manager, Technical Operations

Spark Advisors 11-50 Insurance

Spark Advisors is hiring a Technical Operations leader to own the IT, security, DevOps, and infrastructure function for its AWS-native, regulated Medicare platform as it scales.

AWS CI/CD CloudFormation Encryption HIPAA Secrets Management Terraform
4 hours, 21 minutes ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers