Mozn

Mozn

MOZN is an enterprise AI company that has helped 100+ organizations make critical and informed decisions through specialized AI, in two key areas: Financial Crime Prevention and Enterprise Knowledge Intelligence

Internet Software & Services
51-250

Description

  • Design, deploy, and operate enterprise AI/ML platforms and self-service tooling for data scientists and ML engineers.
  • Deploy and manage Kubeflow, MLflow, KServe, Ray, or similar AI platforms.
  • Design infrastructure for model training, experimentation, feature engineering, and inference.
  • Build highly available and scalable model-serving infrastructure.
  • Design and operate GPU clusters for large-scale AI workloads.
  • Optimize GPU scheduling, utilization, sharing, autoscaling, resource allocation, and distributed training performance.
  • Build CI/CD pipelines and automate AI infrastructure provisioning with Infrastructure as Code.
  • Implement monitoring and observability for GPU utilization, training jobs, model serving, and inference latency.
  • Troubleshoot infrastructure performance bottlenecks and collaborate with data science teams to improve platform usability, reliability, and performance.

Requirements

  • 4–6 years of experience in AI Infrastructure, MLOps, Platform Engineering, or Cloud Engineering.
  • Strong hands-on Kubernetes experience.
  • Experience with Kubeflow, MLflow, or similar ML platforms.
  • Experience operating GPU infrastructure, including NVIDIA technologies, CUDA fundamentals, and GPU optimization.
  • Experience supporting distributed training workloads.
  • Experience with model-serving platforms such as KServe, Triton Inference Server, Ray Serve, or similar.
  • Experience with AWS, GCP, OCI, or Azure AI platforms.
  • Experience with Terraform, Helm, GitOps, or Ansible for infrastructure automation.
  • Strong scripting or programming skills in Python, Bash, or Go.
  • Experience with Prometheus, Grafana, OpenTelemetry, ELK/OpenSearch, or equivalent observability tools.
  • Experience with PyTorch, TensorFlow, Hugging Face, or JAX is preferred.
  • Experience with Ray, DeepSpeed, Horovod, NCCL, vector databases, LLM infrastructure, RAG, or GenAI platforms is preferred.
  • Experience operating LLM inference platforms or supporting AI research and data science teams in production is preferred.
  • Cloud, Kubernetes, NVIDIA, or AI/ML certifications and open-source contributions are advantageous.

Benefits

  • Competitive compensation and top-tier health insurance.
  • High ownership, autonomy, and trust in a dynamic work environment.
  • Opportunity to work with leading AI professionals on high-impact projects.
  • Collaborative and inclusive workplace that values diversity and individual growth.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Machine Learning Engineer

Air 11-50 Media

Air is seeking a Senior Machine Learning Engineer to build and operate the infrastructure supporting the full lifecycle of language models and AI systems for government and industrial customers, from experimentation through production deployment and continuous improvement.

AWS Azure Deep Learning GCP Kubernetes LLM Machine Learning Python Reinforcement Learning
5 hours, 19 minutes ago

RL Environment Software Engineer

Talentpluto, Inc. 1-10 Recruiting

As an RL Environment Software Engineer at a remote applied AI research lab, you will build scalable reinforcement-learning environments and agents that simulate real-world workflows for leading AI research organizations.

Docker Kubernetes Machine Learning Node.js Python React Reinforcement Learning TypeScript
1 day, 5 hours ago

AI Workflow Engineer/Automation Architect

Modern Family Law 51-250 Specialized Consumer Services

Modern Family Law is hiring an AI Workflow Engineer / Automation Architect to build scalable AI-native legal workflows, integrations, and operational automation systems that improve compliant, efficient, client-centered service delivery.

CI/CD Generative AI Git Python
1 day, 5 hours ago

Staff, Machine Learning Engineer - BEV/Multi-Modal Perception

Torc 251-1K Road & Rail

Torc, a Daimler company developing automated-truck software, is seeking a Staff Machine Learning Engineer to lead BEV and multi-modal perception model innovation for autonomous driving.

Computer Vision Deep Learning Machine Learning MLOps Python PyTorch TensorFlow
2 days, 4 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers