Machine Learning Systems Engineer

4 months, 1 week ago
Full-time
Senior
DevOps and Infrastructure
Motional

Motional

Motional is a leading company in driverless technology and autonomous vehicles, leveraging decades of industry expertise to develop and deploy safe and reliable autonomous vehicles. With a powerful DNA combining Aptiv's automotive technology and Hyunda...

Automotive
1K-5K
Founded 2020
$20M raised

Description

  • Profile and optimize training performance by identifying bottlenecks in data loading, gradient computation, and communication.
  • Implement training optimizations such as kernel fusion, sharding, and tiling to reduce step time.
  • Optimize distributed training pipelines using PyTorch Distributed and related tooling.
  • Design and maintain high-performance GPU kernels in Triton or CUDA for ML workloads.
  • Improve data loading pipelines to maximize training throughput.
  • Work at the intersection of machine learning research and high-performance systems engineering to improve speed, cost, reliability, and throughput.
  • Help scale large distributed model training and reduce time to convergence for next-generation models.

Requirements

  • Bachelor’s, Master’s degree, or PhD in Computer Science, Computer Engineering, or a related technical discipline.
  • Strong proficiency in Python.
  • Extensive hands-on experience with PyTorch.
  • Experience optimizing machine learning model execution during training and inference.
  • Strong understanding of fundamental machine learning concepts, architectures, and processes.
  • Exceptional analytical and problem-solving skills.
  • Bias for action and a data-driven approach to technical challenges.
  • Experience with profiling tools such as Nsight and PyTorch Profiler is preferred.
  • Experience with Triton or CUDA is preferred.
  • Experience with distributed training frameworks such as PyTorch Distributed is preferred.

Benefits

  • Base salary range of $144,000 to $192,000 USD.
  • Additional compensation may include a bonus or company equity.
  • Medical, dental, and vision coverage.
  • 401(k) with company match.
  • Health savings accounts.
  • Life insurance.
  • Pet insurance.
  • Hybrid schedule with in-office time in Boston, Pittsburgh, or Las Vegas, or fully remote work available.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Infrastructure Engineer

Sysdig 251-1K IT Services

As a Site Reliability/Infrastructure Engineer at Sysdig, you will build and operate multi-cloud and on-premise infrastructure while improving the reliability, scalability, security, and performance of production systems.

AWS Azure Bash Docker Go Kubernetes Linux Microservices Python
13 minutes ago

Manager, AI/ML Engineering

Lyra Health 1K-5K Health Care Providers & Services

Lyra Health is seeking an engineering leader to execute its machine learning roadmap by scaling and mentoring an AI/ML engineering team that delivers reliable, production-grade systems for mental health care.

AWS HIPAA Java Kotlin Kubernetes Machine Learning Microservices MLOps Neural Networks Prototyping Transformers
23 hours, 58 minutes ago

Senior/Staff Machine Learning Engineer (Model Dev)

Artera 51-250 Construction & Engineering

Artera is seeking an experienced machine learning engineer to lead end-to-end AI biomarker development for cancer care, from clinical problem definition and model validation through regulatory submission and production deployment.

AWS Deep Learning Kubernetes Machine Learning PyTorch TensorFlow
1 day ago

Cloud Security Engineer

Smile Digital Health 251-1K IT Services

Smile Digital Health is seeking a Cloud Security Engineer to design, automate, deploy, and support secure production-grade infrastructure for healthcare data platforms across AWS, Azure, OCI, and GCP.

Ansible AWS Azure Docker HIPAA Kubernetes OpenShift Terraform
1 day, 23 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers