Boundless

Boundless

Boundless provides a protocol for accessing verifiable compute across various blockchain networks, enabling rapid upgrades to zero-knowledge rollups while ensuring transparency and security in trade settlements.

Internet Software & Services
Founded 2005

Description

  • Own AI features and products from prototype through production, including model selection, serving, evaluation, and iteration.
  • Deploy and optimize LLM inference across the GPU fleet using vLLM and SGLang.
  • Tune serving systems for throughput, latency, and cost using techniques such as continuous batching, KV-cache management, quantization, speculative decoding, and multi-model routing.
  • Build and operate reinforcement-learning and post-training pipelines using slime and Prime Intellect tooling.
  • Design reward functions and verifiers, orchestrate rollouts, synchronize weights, and keep long-running training jobs stable.
  • Build evaluation harnesses and benchmarks that measure quality, throughput, and cost together.
  • Use evaluation results and customer/internal feedback to drive fast iteration on products and infrastructure.
  • Partner with Infrastructure on GPU scheduling and fleet utilization.
  • Collaborate with Product on what to build next and why.

Requirements

  • 3+ years shipping ML/AI systems to production.
  • Hands-on experience serving LLM inference with vLLM, SGLang, or TensorRT-LLM.
  • Experience with RL/post-training methods such as GRPO, PPO, DPO, or SFT, or strong adjacent experience with a clear desire to go deep.
  • Strong Python and PyTorch skills.
  • Working understanding of GPU execution, including batching, memory, and basic CUDA concepts.
  • Comfort operating in ambiguity with a strong bias for action.
  • Public GitHub profile required in the application.
  • At least 1 year of GitHub activity/history is required.
  • Direct experience with slime, prime-rl, the verifiers library, or Megatron-LM is preferred.
  • Distributed training experience with FSDP or TP/PP/DP parallelism is preferred.
  • Experience with quantization (FP8/INT8), P/D disaggregation, or speculative decoding is preferred.
  • Experience with verifiable inference or large-scale distributed systems is preferred.
  • Kubernetes and container-based deployment experience is preferred.
  • Familiarity with GPU fleet orchestration tools such as Ray, SkyPilot, or Slurm is preferred.

Benefits

  • Competitive salary of US$175k-$250k annually plus equity allocation.
  • Health, dental, and vision coverage for U.S. employees, with region-adjusted coverage globally.
  • Flexible PTO.
  • Professional development and conference travel budget.
  • Remote-first work with regular off-sites.
  • High-trust, high-velocity team environment.
  • Global hiring, with applicants from around the world welcome.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

AI/ML Engineer

Fieldwire 251-1K Construction & Engineering

Fieldwire, a Hilti construction technology company, is hiring a fully remote AI/ML Engineer in the United States to build and deploy machine learning systems that convert 360° jobsite captures into actionable insights for construction teams.

Agile Computer Vision Deep Learning Generative AI Machine Learning
16 hours, 37 minutes ago

AI Product Engineer

Kyivstar 1K-5K Wireless Telecommunication Services

Kyivstar.Tech is seeking an AI Product Engineer to turn business needs into AI-powered products by combining product ownership, full-stack development, and production delivery.

AWS Azure CI/CD Cybersecurity Docker GraphQL MLOps Node.js Python React REST API
1 day, 17 hours ago

Senior Machine Learning Engineer

Coderio 51-250 Internet Software & Services

Coderio is seeking a Senior Machine Learning Ops Engineer to architect, automate, and deploy scalable machine learning solutions while connecting data science with software engineering across international client initiatives.

Apache Airflow Apache Spark AWS Bash Docker FastAPI Git Machine Learning MLOps Python SQL
1 day, 17 hours ago

Senior ML Engineer (Europe-based/Remote)

SWORD Health 251-1K Health Care Providers & Services

Join Sword’s Europe-based AI team to build and operate reliable, AI-native healthcare systems that improve clinical care at scale.

Machine Learning
1 day, 17 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers