CUDA Engineer

3 weeks, 1 day ago
Full-time
Senior
Software Development
Fuse Energy

Fuse Energy

Fuse Energy is a leading UK electricity supplier prioritizing affordability, service excellence, and sustainability through renewable energy projects and global reinvestment efforts.

Renewable Electricity
11-50
Founded 2014
$78M raised

Description

  • Write and optimize custom CUDA kernels for core transformer inference operations.
  • Profile kernels to identify and eliminate bottlenecks in occupancy, memory throughput, and warp divergence.
  • Apply kernel fusion to reduce memory round-trips and launch overhead across inference pipelines.
  • Optimize memory access patterns and manage the memory hierarchy for maximum bandwidth utilization.
  • Implement quantisation-aware kernels and mixed-precision arithmetic to reduce latency and memory footprint.
  • Build and tune caching mechanisms for efficient autoregressive decoding.
  • Tune kernel launch configurations for target GPU architectures.
  • Benchmark kernels against existing baselines and drive measurable throughput and latency improvements.
  • Write tests for CUDA code to catch performance and correctness regressions.
  • Maintain internal CUDA libraries and contribute to coding standards and documentation.

Requirements

  • 4+ years writing production CUDA code, with a track record of shipping performance-critical kernels.
  • Deep understanding of GPU microarchitecture, warps, occupancy, register pressure, and memory hierarchy.
  • Strong CUDA C++ skills, including streams and asynchronous execution.
  • Hands-on experience profiling to diagnose compute-bound vs. memory-bound bottlenecks.
  • Experience with kernel fusion, memory coalescing, and avoiding warp divergence.
  • Experience writing quantised and mixed-precision kernels.
  • Solid grasp of parallel algorithm design and numerical precision tradeoffs.
  • Experience with transformer/attention-style kernels or autoregressive decoding (nice to have).
  • Experience building high-performance GPU libraries from scratch (nice to have).
  • Background in HPC or other latency-critical performance engineering (nice to have).
  • Exposure to multi-GPU or multi-node kernel-level optimisation (nice to have).
  • Comfortable reading PTX/SASS to validate kernel efficiency (nice to have).

Benefits

  • Competitive salary and an equity sign-on bonus.
  • Biannual bonus scheme.
  • Fully expensed tech to match your needs.
  • Breakfast and dinner allowance for office-based employees.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

AI Developer

AI Developer (Contract) at a remote company, responsible for building AI-powered workflows and internal tools that improve business efficiency and create new capabilities.

JavaScript LLM Python
16 hours, 46 minutes ago

Staff ML Engineer (ML/AI)

Lyra Health 1K-5K Health Care Providers & Services

Lyra Health is hiring a Staff ML/AI Engineer to define the architecture and technical strategy for enterprise-scale machine learning and generative AI platforms supporting clinically precise, reliable mental health products.

AWS Celery CI/CD Docker Generative AI HIPAA Java Kafka Kotlin Kubernetes Machine Learning Microservices MLOps Python REST API
1 day, 16 hours ago

AI Agent Implementations Specialist

Extreme Networks 1K-5K IT Services

Extreme Networks is hiring an AI Agent Developer & Marketo Administrator to support marketing automation, AI agent implementation, and MarTech operations across its global marketing organization.

CRM JavaScript Machine Learning Node.js Python REST API
1 day, 16 hours ago

Entrepreneur in Residence - Technical Co-founder (CTO)

FutureSight 11-50 Internet Software & Services

FutureSight is seeking a Co-Founder & CTO to lead the technical direction of a new B2B AI venture from inception to launch and early growth.

LLM System Design
1 day, 16 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers