Mozn

Mozn

MOZN is an enterprise AI company that has helped 100+ organizations make critical and informed decisions through specialized AI, in two key areas: Financial Crime Prevention and Enterprise Knowledge Intelligence

Internet Software & Services
51-250

Description

  • Design, deploy, and operate enterprise AI/ML platforms and self-service tooling for data scientists and ML engineers.
  • Deploy and manage Kubeflow, MLflow, KServe, Ray, or similar AI platforms.
  • Design infrastructure for model training, experimentation, feature engineering, and inference.
  • Build highly available and scalable model-serving infrastructure.
  • Design and operate GPU clusters for large-scale AI workloads.
  • Optimize GPU scheduling, utilization, sharing, autoscaling, resource allocation, and distributed training performance.
  • Build CI/CD pipelines and automate AI infrastructure provisioning with Infrastructure as Code.
  • Implement monitoring and observability for GPU utilization, training jobs, model serving, and inference latency.
  • Troubleshoot infrastructure performance bottlenecks and collaborate with data science teams to improve platform usability, reliability, and performance.

Requirements

  • 4–6 years of experience in AI Infrastructure, MLOps, Platform Engineering, or Cloud Engineering.
  • Strong hands-on Kubernetes experience.
  • Experience with Kubeflow, MLflow, or similar ML platforms.
  • Experience operating GPU infrastructure, including NVIDIA technologies, CUDA fundamentals, and GPU optimization.
  • Experience supporting distributed training workloads.
  • Experience with model-serving platforms such as KServe, Triton Inference Server, Ray Serve, or similar.
  • Experience with AWS, GCP, OCI, or Azure AI platforms.
  • Experience with Terraform, Helm, GitOps, or Ansible for infrastructure automation.
  • Strong scripting or programming skills in Python, Bash, or Go.
  • Experience with Prometheus, Grafana, OpenTelemetry, ELK/OpenSearch, or equivalent observability tools.
  • Experience with PyTorch, TensorFlow, Hugging Face, or JAX is preferred.
  • Experience with Ray, DeepSpeed, Horovod, NCCL, vector databases, LLM infrastructure, RAG, or GenAI platforms is preferred.
  • Experience operating LLM inference platforms or supporting AI research and data science teams in production is preferred.
  • Cloud, Kubernetes, NVIDIA, or AI/ML certifications and open-source contributions are advantageous.

Benefits

  • Competitive compensation and top-tier health insurance.
  • High ownership, autonomy, and trust in a dynamic work environment.
  • Opportunity to work with leading AI professionals on high-impact projects.
  • Collaborative and inclusive workplace that values diversity and individual growth.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Staff Machine Learning Engineer

Babylist 251-1K Internet Software & Services

Babylist is hiring a Staff Machine Learning Engineer to define and build personalization systems across its registry, commerce, health, search, and AI products used by millions of families.

Apache Airflow AWS CI/CD CRM Datadog dbt Deep Learning GitHub Actions Kotlin Kubernetes Linear Machine Learning MLflow MySQL Pandas Python PyTorch React Ruby on Rails SageMaker Scikit-learn Shopify Sidekiq Snowflake Swift System Design Terraform TypeScript XGBoost
2 hours, 35 minutes ago

Staff Machine Learning Engineer

Babylist 251-1K Internet Software & Services

Babylist is hiring a Staff Machine Learning Engineer to lead personalization across feeds, recommendations, search, and AI products used by millions of families.

Apache Airflow AWS CI/CD CRM Datadog dbt Deep Learning GitHub Actions Kotlin Kubernetes Linear Machine Learning MLflow MLOps MySQL Pandas Python PyTorch React Ruby Ruby on Rails Scikit-learn Shopify Sidekiq Snowflake Swift System Design Terraform TypeScript XGBoost
2 hours, 50 minutes ago

Staff Machine Learning Engineer, Ads Creative Effectiveness

Reddit 1K-5K Internet Software & Services

Reddit is hiring a Staff engineer for its Ads Creative Effectiveness team to build GenAI and predictive products that generate, edit, optimize, and safeguard high-performing ad creative.

Generative AI Machine Learning Neural Networks
2 hours, 50 minutes ago

Junior AI/ML Engineer (GenAI, AWS)

Provectus 251-1K Professional Services

Provectus is hiring a software/ML engineer to build and deploy applied AI systems, including RAG applications and agentic solutions, for enterprise clients in financial services, insurance, healthcare, and life sciences.

Apache Airflow Apache Spark AWS CI/CD Go Kafka Microservices NLP Python REST API Rust TypeScript
3 hours, 5 minutes ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers