Tech Lead Manager ML Optimization

3 months, 3 weeks ago
Full-time
Lead
Software Development

Waymo

Waymo is an autonomous driving technology company building the Waymo Driver and operating Waymo One, its fully autonomous ride-hailing service.

Autonomous vehicles, robotics, AI, ride-hailing / mobility tech
Founded 2009
$21600M raised

Description

  • Lead the development and enable efficient deployment of large-scale machine learning models using advanced AI infrastructure.
  • Own model efficiency improvements across multiple platforms and drive model-system co-design to meet technical and business requirements.
  • Study state-of-the-art model architectures and optimizations and translate them into measurable deliverables for Waymo’s driving stack.
  • Build developer tooling for model performance inspection in distributed training and inference environments.
  • Apply roofline analysis to identify efficiency gaps and drive optimization work to meet system requirements.
  • Innovate high-performance optimization techniques and tools for large-scale training and inference, including future TPU and low-bit precision setups.
  • Guide cross-team efforts across data generation, model development, and deployment pipelines.
  • Mentor junior engineers and foster collaboration and engineering excellence.
  • Manage individual contributor performance for a medium-sized team of about 10 engineers.
  • Work at the intersection of data engineering, model development, and low-latency datacenter and on-device deployment.

Requirements

  • 10+ years of professional software engineering experience.
  • At least 5 years of experience in machine learning infrastructure, including developing, training, deploying, and optimizing large-scale ML systems.
  • Experience using ML accelerator profiling tools to identify performance bottlenecks.
  • Experience with ML infrastructure tools or frameworks such as DeepSpeed, PyTorch, TensorFlow, JAX, or similar.
  • Deep understanding of modern ML models and architectures, including autoregressive and diffusion transformers.
  • Familiarity with custom kernels for efficiency across diverse hardware compute platforms.
  • Strong leadership experience working across cross-functional teams and multiple organizations.
  • Excellent verbal and written communication skills with the ability to explain complex technical concepts to broad audiences.
  • A Master’s or PhD in Computer Science, Engineering, or a related field is preferred.

Benefits

  • Base salary range of $298,000 to $378,000 USD.
  • Eligibility for Waymo’s discretionary annual bonus program.
  • Eligibility for Waymo’s equity incentive plan.
  • Access to Waymo’s generous company benefits program, subject to eligibility requirements.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Engineering Manager – Unified Telephony Platform

Eltropy 51-250 Communications Equipment

Remote Engineering Manager role at a communications platform company leading the buildout of a cloud-native unified telephony platform for financial institutions.

CI/CD Go Microservices Python Twilio
15 hours, 59 minutes ago

Staff ML Engineer (ML/AI)

Lyra Health 1K-5K Health Care Providers & Services

Lyra Health is hiring a Staff ML/AI Engineer to define the architecture and technical strategy for enterprise-scale machine learning and generative AI platforms supporting clinically precise, reliable mental health products.

AWS Celery CI/CD Docker Generative AI HIPAA Java Kafka Kotlin Kubernetes Machine Learning Microservices MLOps Python REST API
1 day, 16 hours ago

Applied AI/ML Engineer

Boundless Internet Software & Services

Boundless is hiring an Applied AI/ML Engineer to build and ship end-to-end AI products on its GPU inference fleet, from low-latency serving through reinforcement-learning post-training and production evaluation.

Docker Kubernetes Python PyTorch Reinforcement Learning
1 day, 16 hours ago

Sr. ML Engineer (MLOps)

Lyra Health 1K-5K Health Care Providers & Services

Lyra Health is hiring a Senior ML Engineer to build and operate the tooling and services behind production machine learning and generative AI products for its mental health platform.

AWS Celery Docker Generative AI Java Kotlin Kubernetes Machine Learning MLOps Python REST API
1 day, 16 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers