Nebius

Nebius

Nebius enables B2B companies to build local hyperscaling cloud platforms with cost-effective GPUs, InfiniBand network, and 50% less compute cost. They offer managed Kubernetes and a launch-ready business model for innovative cloud solutions.

Internet Software & Services
51-250

Description

  • Work with hardware and development teams to profile and analyze GPU performance at the system and kernel level.
  • Evaluate and compare GPU performance across different platforms, architectures, and software stacks such as CUDA and ROCm.
  • Debug and optimize ML workloads to run efficiently on GPU hardware by identifying and resolving performance bottlenecks.
  • Perform acceptance testing for new GPU clusters to verify performance, stability, and compatibility for AI workloads.
  • Run experiments across diverse GPU system configurations to assess the impact of interconnect strategies and system-level optimizations.
  • Develop tools and dashboards to visualize performance metrics, bottlenecks, and trends.
  • Contribute to internal tooling, frameworks, and best practices for GPU benchmarking and optimization.

Requirements

  • A strong understanding of the theoretical foundations of machine learning.
  • Deep understanding of performance aspects of large neural network training and inference, including data, tensor, context, and expert parallelism, offloading, custom kernels, hardware features, attention optimizations, and dynamic batching.
  • Deep experience with modern deep learning frameworks such as PyTorch, JAX, Megatron-LM, and TensorRT-LLM.
  • Good understanding of the GPU stack, including CUDA, NCCL, drivers, and relevant libraries.
  • Familiarity with containerized environments such as Docker and Kubernetes.
  • Strong communication skills and the ability to work independently.
  • Familiarity with modern LLM inference frameworks such as vLLM, SGLang, and TensorRT, preferred.
  • Experience with Python and performance profiling tools such as Nsight, nvprof, and perf, preferred.
  • Familiarity with cloud ML platforms such as AWS, GCP, and Azure ML, preferred.
  • Contributions to open-source ML benchmarking tools, preferred.
  • Authorization to work in the country of application, with proof of employment eligibility required at hire.

Benefits

  • Competitive compensation.
  • Career growth and learning opportunities.
  • Flexibility and work-life balance.
  • Collaborative and innovative culture.
  • Opportunity to work on impactful AI projects.
  • International environment and talented teams.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Storage Engineer

Ensono 1K-5K IT Services

Ensono is seeking a Storage Engineer to support the stability, maintenance, and optimization of storage area networks, NAS arrays, and array-based replication across managed environments.

22 hours, 24 minutes ago

Staff Machine Learning Engineer, Applied Research

Pinterest 5K-10K Internet Software & Services

Pinterest Labs is hiring a research-focused machine learning role to advance LLMs and applied AI work across search, recommendations, and other core Pinterest systems.

C++ Computer Vision Java LLM Machine Learning MLflow NLP Python PyTorch Reinforcement Learning TensorFlow
22 hours, 54 minutes ago

Senior ML Engineer (LLMs, AWS)

Provectus 251-1K Professional Services

Provectus is hiring an ML Engineer to develop and maintain production machine learning solutions across AI, cloud, and data engineering projects.

Apache Spark AWS Deep Learning Docker Feature Engineering LLM Machine Learning MLOps NLP Python
22 hours, 54 minutes ago

[Job-30851] Master Software Developer/Devops (NVIDIA knowledge), Colombia

CI&T 5K-10K Internet Software & Services

CI&T is hiring a technical leader to develop and operate GPU-accelerated routing and optimization solutions on NVIDIA cuOpt and Microsoft Azure for large-scale enterprise operations.

Azure CI/CD
23 hours, 9 minutes ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers