Lead Machine Learning Engineer, Inference & Performance

2 weeks, 6 days ago
Full-time
Senior
Software Development
Egen.ai

Egen.ai

Egen.ai specializes in providing technology services that leverage cloud computing, data analytics, and artificial intelligence to enhance document intelligence and drive productivity and growth for its clients.

IT Services
Founded 2000

Description

  • Own the full lifecycle of AI features from initial prototype to robust, scalable production services.
  • Design and optimize production LLM serving to maximize throughput and minimize latency.
  • Instrument and profile training runs to identify bottlenecks and improve performance.
  • Tune attention implementations and other inference/training techniques for specific hardware platforms.
  • Deploy and operate multiple models within shared GPU clusters on Google Kubernetes Engine (GKE).
  • Improve GPU utilization, throughput-per-dollar, and overall fleet efficiency.
  • Collaborate with clients to translate business needs and constraints into AI architectures.
  • Write clean, maintainable code and apply a disciplined software engineering approach to AI systems.

Requirements

  • Bachelor's or Master's degree in Computer Science, Engineering, or a related technical field.
  • 5+ years of experience in ML/AI engineering, with meaningful experience in performance, infrastructure, or systems.
  • Proven track record of deploying and optimizing models in a production environment.
  • Demonstrated experience profiling and improving GPU utilization for training and/or inference.
  • Hands-on experience with vLLM, SGLang, or comparable high-performance serving stacks.
  • Strong Kubernetes experience, specifically deploying and autoscaling multiple models on shared GPU clusters on Google Cloud/GKE.
  • Mastery of Python and shell scripting.
  • Solid grasp of GPU architecture, LLM inference fundamentals, and the attention mechanism.
  • Fluency with profiling tools for diagnosing compute-bound and memory-bound bottlenecks.
  • Knowledge of data engineering and SQL.
  • Experience with classic machine learning, neural nets, training, and tuning is a strong plus.
  • Comfort reading lower-level CUDA-adjacent performance code is a strong plus.

Benefits

  • Competitive salary.
  • Comprehensive health insurance.
  • Paid leave, including vacation/PTO.
  • Paid holidays.
  • Sick leave.
  • Parental leave.
  • Bereavement leave.
  • 401(k) employer match.
  • Employee referral bonuses.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

ML Infrastructure Engineer

x.ai 51-250 Internet Software & Services

SpaceXAI is hiring an ML Infrastructure Engineer to build and optimize the machine learning platform that powers recommendations on X.

Ansible C++ Linux Puppet Python PyTorch Rust
1 day, 19 hours ago

Senior Machine Learning Engineer, Safety

Reddit 1K-5K Internet Software & Services

Reddit is hiring a remote-friendly Machine Learning Engineer to build and improve safety systems that support enforcement of Reddit rules using large language models.

Deep Learning LLM Machine Learning NLP Python PyTorch TensorFlow
1 day, 20 hours ago

Senior, Machine Learning Engineer - 3D Perception

Torc 251-1K Road & Rail

Torc is hiring a Senior Machine Learning Engineer – 3D Perception to develop and deploy production perception models for autonomous trucks and improve Bird's Eye View understanding across its autonomy stack.

C++ Computer Vision Deep Learning Machine Learning Python PyTorch
1 day, 20 hours ago

Applied AI Engineer

Unframe Inc. 51-200 Technology, Information and Internet

Unframe is hiring an Applied AI Engineer to work with enterprise customers and internal teams on deploying AI-driven solutions that connect business problems to production systems.

Machine Learning Python React TypeScript
2 days, 18 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers