Applied AI/ML Engineer

14 hours, 35 minutes ago
Full-time
Mid Level
Software Development
Boundless

Boundless

Boundless provides a protocol for accessing verifiable compute across various blockchain networks, enabling rapid upgrades to zero-knowledge rollups while ensuring transparency and security in trade settlements.

Internet Software & Services
Founded 2005

Description

  • Own AI features and products from prototype through production, including model selection, serving, evaluation, and iteration.
  • Deploy and optimize LLM inference across the GPU fleet using vLLM and SGLang.
  • Tune serving systems for throughput, latency, and cost using techniques such as continuous batching, KV-cache management, quantization, speculative decoding, and multi-model routing.
  • Build and operate reinforcement-learning and post-training pipelines using slime and Prime Intellect tooling.
  • Design reward functions and verifiers, orchestrate rollouts, synchronize weights, and keep long-running training jobs stable.
  • Build evaluation harnesses and benchmarks that measure quality, throughput, and cost together.
  • Use evaluation results and customer/internal feedback to drive fast iteration on products and infrastructure.
  • Partner with Infrastructure on GPU scheduling and fleet utilization.
  • Collaborate with Product on what to build next and why.

Requirements

  • 3+ years shipping ML/AI systems to production.
  • Hands-on experience serving LLM inference with vLLM, SGLang, or TensorRT-LLM.
  • Experience with RL/post-training methods such as GRPO, PPO, DPO, or SFT, or strong adjacent experience with a clear desire to go deep.
  • Strong Python and PyTorch skills.
  • Working understanding of GPU execution, including batching, memory, and basic CUDA concepts.
  • Comfort operating in ambiguity with a strong bias for action.
  • Public GitHub profile required in the application.
  • At least 1 year of GitHub activity/history is required.
  • Direct experience with slime, prime-rl, the verifiers library, or Megatron-LM is preferred.
  • Distributed training experience with FSDP or TP/PP/DP parallelism is preferred.
  • Experience with quantization (FP8/INT8), P/D disaggregation, or speculative decoding is preferred.
  • Experience with verifiable inference or large-scale distributed systems is preferred.
  • Kubernetes and container-based deployment experience is preferred.
  • Familiarity with GPU fleet orchestration tools such as Ray, SkyPilot, or Slurm is preferred.

Benefits

  • Competitive salary of US$175k-$250k annually plus equity allocation.
  • Health, dental, and vision coverage for U.S. employees, with region-adjusted coverage globally.
  • Flexible PTO.
  • Professional development and conference travel budget.
  • Remote-first work with regular off-sites.
  • High-trust, high-velocity team environment.
  • Global hiring, with applicants from around the world welcome.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Staff ML Engineer (ML/AI)

Lyra Health 1K-5K Health Care Providers & Services

Lyra Health is hiring a Staff ML/AI Engineer to define the architecture and technical strategy for enterprise-scale machine learning and generative AI platforms supporting clinically precise, reliable mental health products.

AWS Celery CI/CD Docker Generative AI HIPAA Java Kafka Kotlin Kubernetes Machine Learning Microservices MLOps Python REST API
14 hours, 35 minutes ago

Sr. ML Engineer (MLOps)

Lyra Health 1K-5K Health Care Providers & Services

Lyra Health is hiring a Senior ML Engineer to build and operate the tooling and services behind production machine learning and generative AI products for its mental health platform.

AWS Celery Docker Generative AI Java Kotlin Kubernetes Machine Learning MLOps Python REST API
14 hours, 35 minutes ago

Data & ML Engineer

DEFCON AI 11-50 Internet Software & Services

Defcon AI is hiring a Data & ML Engineer to build the data, matching, scoring, and retrieval layers of an AI-enabled decision-support system in a controlled government cloud environment.

Apache Airflow Apache Spark dbt DevSecOps Kafka Machine Learning PostgreSQL Python PyTorch Scikit-learn SQL XGBoost
1 day, 13 hours ago

ML Engineer, II - Simulation Enablement

Torc 251-1K Road & Rail

Torc is hiring a Machine Learning Engineer to support Torc Sim, its autonomous trucking simulation platform, by partnering with autonomy teams to scale replay, recompute, evaluation, and visualization workflows.

GitHub Actions Machine Learning OpenGL Pandas Python Terraform Three.js
1 day, 13 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers