AI Inference Engineer

2 weeks, 3 days ago
Full-time
Senior
DevOps and Infrastructure
Fuse Energy

Fuse Energy

Fuse Energy is a leading UK electricity supplier prioritizing affordability, service excellence, and sustainability through renewable energy projects and global reinvestment efforts.

Renewable Electricity
11-50
Founded 2014
$78M raised

Description

  • Define Fuse's inference serving strategy and architecture from first principles.
  • Design and build the serving stack for request routing, batching, scheduling, and autoscaling.
  • Own model-level optimization for serving, including quantization, distillation, and speculative decoding.
  • Select and evaluate serving frameworks and orchestration tools such as vLLM, TensorRT-LLM, SGLang, and Triton Inference Server.
  • Translate throughput, latency, and uptime commitments into technical specifications and capacity plans.
  • Own inference performance and reliability for the serving layer.
  • Collaborate with CUDA and GPU engineering teams to integrate low-level performance improvements.
  • Set standards, tooling, and benchmarks for the function as it scales.

Requirements

  • 4+ years of experience building or operating large-scale inference serving systems, or equivalent strong project/industry experience.
  • Deep hands-on experience with inference serving frameworks and optimization techniques such as batching, KV-cache management, quantization, and speculative decoding.
  • Strong systems thinking across the full request-to-response path in a large cluster.
  • Comfort working directly with GPU/CUDA engineers to integrate low-level performance work.
  • Track record of making high-stakes architecture decisions and owning the outcome.
  • Ability to operate without a playbook in a founding, early-stage role.
  • Experience with Triton or custom ML inference/training frameworks (preferred).
  • Experience with autoscaling or capacity planning for large-scale inference workloads (preferred).
  • Exposure to multi-tenant serving or SLA-driven infrastructure (preferred).
  • Background at a hyperscaler, frontier AI lab, or large-scale distributed inference system (preferred).
  • Familiarity with Kubernetes or Slurm for cluster orchestration (preferred).
  • Interest or experience in energy markets, grid systems, or sustainability-focused compute (preferred).

Benefits

  • Competitive salary and an equity sign-on bonus.
  • Biannual bonus scheme.
  • Fully expensed tech to match your needs.
  • Breakfast and dinner allowance for office-based employees.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Data Engineer + AI

Coforge 10K-50K IT Services

Coforge is hiring a Senior Data Engineer + AI to build and maintain cloud-native data platforms and AI-powered applications for internal and customer-facing use.

Apache Spark AWS CI/CD Databricks dbt LLM Python SQL TypeScript
14 hours, 31 minutes ago

Applied AI solutions Architect

phData 251-1K IT Services

phData is hiring an Applied AI Solutions Architect to design and deliver production-ready AI solutions for enterprise clients, bridging data, cloud, and AI systems to create measurable business impact.

AWS Azure CI/CD Databricks dbt GCP Generative AI MLOps Python SageMaker Snowflake SQL
1 day, 13 hours ago

Senior AI Product Builder, Inventory Optimization

ShipBob 251-1K Air Freight & Logistics

ShipBob is hiring a remote Senior AI Product Builder to own the end-to-end build and optimization of its inventory decision systems across the fulfillment network.

Microservices Prototyping Python REST API SQL TypeScript
1 day, 13 hours ago

Perception Engineer

Motional 1K-5K Automotive

Motional is hiring a Perception Engineer in Boston or remote U.S. to develop and improve AI perception systems for autonomous vehicles.

C++ Git Machine Learning Python PyTorch
1 day, 13 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers