Principal Machine Learning Engineer

1 month, 3 weeks ago
Full-time
Lead
Software Development
AVOMIND

AVOMIND

AVOMIND is a global recruitment firm based in Berlin, specializing in hiring commercial, strategy, and analytics & insights talent for companies and candidates worldwide. With a focus on executive search and embedded recruitment, AVOMIND connects high-...

Professional Services
11-50
Founded 2019

Description

  • Build and own end-to-end machine learning pipelines from data processing through training, evaluation, inference, and deployment.
  • Fine-tune and adapt models using modern techniques such as LoRA, QLoRA, SFT, DPO, and model distillation.
  • Design and operate scalable inference systems while balancing latency, cost, and reliability.
  • Develop and maintain data pipelines for synthetic and real-world training datasets.
  • Build evaluation frameworks to measure model performance, robustness, safety, and bias.
  • Optimize production deployments through GPU efficiency, memory usage, latency reduction, and scaling strategies.
  • Collaborate with application engineering teams to integrate ML systems into backend, mobile, and desktop applications.
  • Monitor production ML systems and iterate quickly based on real-world performance and constraints.

Requirements

  • Strong background in deep learning and transformer-based architectures.
  • Hands-on experience training, fine-tuning, or deploying large-scale machine learning models in production.
  • Proficiency with machine learning frameworks such as PyTorch or JAX.
  • Experience with distributed training and inference frameworks such as DeepSpeed, FSDP, Megatron, ZeRO, or Ray.
  • Strong software engineering skills and experience building robust, maintainable, production-grade systems.
  • Experience optimizing GPU workloads, including memory efficiency, quantization, and mixed precision.
  • Ability to independently own end-to-end machine learning systems in fast-moving environments.
  • Strong problem-solving skills with a focus on rapid iteration and continuous improvement.
  • Experience with LLM inference frameworks such as vLLM, TensorRT-LLM, or FasterTransformer (preferred).
  • Open-source contributions to machine learning or systems libraries (preferred).
  • Scientific computing, compiler technologies, or GPU kernel development (preferred).
  • Experience with RLHF pipelines, including PPO, DPO, or ORPO (preferred).
  • Experience training or deploying multimodal or diffusion models (preferred).
  • Experience with large-scale data processing frameworks such as Apache Arrow, Spark, or Ray (preferred).

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Lead DataOps / MLOps

Protective Life 1K-5K Insurance

Protective Life is seeking a hands-on DataOps/MLOps Lead for its Voyager product pod to build and operate reliable, secure, governed data and machine-learning platforms on Databricks Lakehouse and Microsoft Azure in a regulated insurance environment.

Azure CI/CD Dagster Databricks dbt Docker Git Kubernetes MLflow MLOps Python Secrets Management SQL Terraform
3 hours, 48 minutes ago

Senior Machine Learning Systems Engineer (US)

PointClickCare 1K-5K Health Care Providers & Services

PointClickCare is seeking a Senior Machine Learning Systems Engineer to build and operate the scalable ML platform, infrastructure, and tooling that enables teams to develop, deploy, and monitor AI solutions in healthcare.

AWS Azure CI/CD Databricks Docker Java Kubeflow Kubernetes Machine Learning MLflow MLOps Network Security Python
3 hours, 48 minutes ago

Staff Engineer, AI Operations & Governance, Workplace AI (R5428)

Bitly 51-250 Internet Software & Services

Shield AI is hiring a Staff Engineer for Enterprise AI Operations & Governance to operate workplace AI platforms, respond to production incidents, and translate technical operations into governance and risk practices.

JSON Secrets Management YAML
4 hours, 3 minutes ago

Data Science ML/Gen AI Engineer (Mid-Level)- Orbit

Irth 51-250 Diversified Telecommunication Services

Irth Solutions is hiring a remote India-based ML/GenAI Engineer to build governed Databricks Lakehouse foundations and production AI solutions that deliver measurable insights across damage prevention, asset integrity, land management, and stakeholder engagement.

Apache Spark AWS Azure CI/CD Databricks DynamoDB Feature Engineering Generative AI GitHub Actions JIRA Machine Learning MLflow Power BI Python SQL
4 hours, 18 minutes ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers