Engineering Lead, Inference Optimization

18 hours, 55 minutes ago
Full-time
Lead
Software Development
Venice.ai

Venice.ai

Try Venice.ai for free. Generate text, images, characters and video using private and unbiased AI.

Founded 2019

Description

  • Own Venice’s technical strategy for inference performance.
  • Recruit and lead the Inference Optimization Team.
  • Optimize GPU infrastructure across multiple architectures, including H200s and B300s.
  • Improve latency, throughput, and cost per token for LLM inference workloads.
  • Build reproducible benchmarking harnesses across inference engines to compare engines, quantization schemes, and parallelism strategies.
  • Work with the inference routing system to improve multivariate load-balancing algorithms.
  • Evaluate new inference optimization methods, including custom CUDA/Triton kernels, attention variants, quantization schemes, and compilation improvements.
  • Assess emerging inference hardware such as FPGAs, ASICs, and custom silicon for fit within Venice’s stack.

Requirements

  • 8+ years of experience in performance optimization or HPC, with deep knowledge of GPU architecture and parallel programming.
  • 5+ years of experience leading engineering teams.
  • Proficiency in Python, Rust, or Go.
  • Bonus: C++/CUDA experience.
  • Hands-on experience with at least one production LLM inference engine such as vLLM or SGLang at high volume.
  • Experience with LLM inference optimization techniques, including continuous batching, PagedAttention/KV cache management, speculative decoding, quantization, CUDA graphs, and torch.compile.
  • Strong understanding of quantization tradeoffs, both qualitative and quantitative.
  • Experience with distributed inference strategies in multi-GPU and multi-node environments, including tensor parallelism, pipeline parallelism, and MoE parallelism.
  • Fluency with GPU profiling tools such as Nsight Systems, Nsight Compute, and PyTorch Profiler, with a bias toward measurement before optimization.
  • Bonus: experience optimizing diffusion/image models, building custom Triton kernels, or contributing to open-source inference frameworks.

Benefits

  • Base annual salary of $270,000-$330,000 USD.
  • Reports to the Head of Engineering.
  • Opportunity to work on privacy-focused AI at the bleeding edge of inference performance.
  • High-impact role with both individual contributor work and team management.
  • Opportunity to shape technical strategy and build a team from the ground up.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Scientific Platform Engineer

VivoSense 11-50 Professional Services

VivoSense is hiring a Senior Scientific Platform Engineer to transform scientific and digital health workflows into scalable, production-ready platform capabilities for clinical studies and customer deliverables.

AWS C# CI/CD .NET Python R Serverless
19 hours, 10 minutes ago

Backend Engineering Manager

Dragos 251-1K Professional Services

Dragos is hiring a Backend Software Engineering Manager to lead the backend team that develops and sustains core platform services for xOT cybersecurity in support of critical infrastructure.

Agile Cybersecurity HashiCorp Vault Microservices RabbitMQ Rust
1 day, 19 hours ago

Senior Engineering Manager

Civica 1K-5K Internet Software & Services

Civica is seeking a Senior Engineering Manager to lead a multi-team engineering domain supporting essential public services and drive delivery, team health, and technical improvement across a mix of legacy and modern systems.

1 day, 19 hours ago

Associate Research Scientist, Real World Evidence

Precision AQ 1001-5000 Business Consulting and Services

Precision AQ is hiring a fully remote Associate Research Scientist on its Real-World Evidence team to support real-world studies in the pharmaceutical and biotech space.

Python R SQL
2 days, 19 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers