AI Platform Support Engineer (APAC)

1 month, 3 weeks ago
Full-time
Senior
Artificial Intelligence and Machine Learning
Lightning AI

Lightning AI

Lightning AI is an all-in-one platform for AI development, enabling users to code, prototype, train, scale, and serve AI models lightning fast. It offers tools to build models, create Lightning Apps, and streamline the ML lifecycle process, allowing fo...

IT Services
11-50
Founded 2019
$59M raised

Description

  • Partner directly with customer engineering teams running training and inference workloads in production.
  • Help customers diagnose and resolve complex distributed systems and ML infrastructure issues.
  • Act as a technical advisor during high-impact incidents and platform degradation events.
  • Investigate failures involving distributed training, Kubernetes orchestration, GPU allocation, networking, and storage systems.
  • Troubleshoot PyTorch, CUDA, NCCL, and inference serving issues.
  • Analyze logs, metrics, traces, and system behavior to identify root causes.
  • Debug containerized workloads running across Kubernetes and bare-metal GPU environments.
  • Identify recurring customer issue patterns and drive long-term reliability improvements.
  • Contribute to post-incident reviews, internal tooling, automation, documentation, and runbooks.
  • Partner with infrastructure, networking, and platform engineering teams to improve observability and troubleshooting workflows.

Requirements

  • Strong software engineering and systems troubleshooting background.
  • Experience with Kubernetes and containerized environments.
  • Linux systems knowledge, including networking, storage, process management, and performance tuning.
  • Experience with cloud infrastructure and distributed systems.
  • Experience with observability and debugging tools such as Prometheus, Grafana, or OpenTelemetry.
  • Hands-on experience operating machine learning workloads in production or research environments.
  • Experience with distributed ML systems and tooling such as PyTorch, CUDA, or NCCL.
  • Familiarity with GPU infrastructure and orchestration.
  • Experience troubleshooting performance, reliability, or scaling issues in ML infrastructure.
  • Strong communication skills and ability to work directly with highly technical customers and engineering teams.
  • Comfortable operating in fast-moving, highly ambiguous environments.
  • Experience with large-scale model training or distributed inference systems (nice to have).
  • Familiarity with Ray, Kubeflow, Slurm, or similar distributed scheduling platforms (nice to have).
  • Experience with InfiniBand, RDMA, or high-performance networking (nice to have).
  • Experience operating bare-metal infrastructure (nice to have).
  • Familiarity with storage systems commonly used in ML environments (nice to have).
  • Experience working at an AI infrastructure, cloud, MLOps, or developer tooling company (nice to have).
  • Experience writing automation, tooling, or scripts in Python or similar languages (nice to have).
  • This role is remote and open to candidates based in the Philippines.
  • This role follows a Sunday-Wednesday shift schedule, with working hours from 7:00 AM to 5:00 PM local time (UTC+8).

Benefits

  • Medical, dental, and vision coverage for employees and eligible dependents.
  • RSUs with meaningful equity in the company.
  • 401(k) matching in the U.S. and pension contributions in the U.K.
  • Unlimited PTO, company holidays, and floating holidays.
  • Two-week company-wide winter break.
  • Paid parental and family leave.
  • Annual learning and development allowance.
  • Wellness and work-from-home stipends.
  • Four weeks of paid sabbatical leave after four years of service.
  • Flexible schedules and a hybrid work model for office-based teams.
  • Complimentary meals at office hubs.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior MLOps Engineer

Point Wild Internet Software & Services

Point Wild is hiring a Senior MLOps Engineer to build and operate GCP infrastructure, pipelines, and tooling that deploy and scale reliable AI systems for cybersecurity products.

Apache Airflow Argo CD CI/CD Deep Learning Docker GCP GitHub Actions Grafana Kubernetes Machine Learning Microservices MLflow MLOps Prometheus Python SQL Terraform Vertex AI
1 hour, 56 minutes ago

Senior MLOps Engineer

Point Wild Internet Software & Services

Point Wild is hiring a Senior MLOps Engineer to build and operate GCP-based infrastructure, pipelines, and tooling that reliably deploy and scale AI models for cybersecurity products.

Apache Airflow Argo CD CI/CD Cybersecurity Deep Learning Docker GCP GitHub Actions Grafana Kubernetes Machine Learning Microservices MLflow MLOps Prometheus Python SQL Terraform Vertex AI
1 hour, 56 minutes ago

Senior MLOps Engineer

Point Wild Internet Software & Services

Point Wild is seeking a Senior MLOps Engineer to build and operate scalable, reliable machine-learning infrastructure on Google Cloud Platform for cybersecurity products that protect customers’ digital identities and personal information.

Apache Airflow Argo CD CI/CD Cybersecurity Deep Learning Docker GCP GitHub Actions Grafana Kubernetes Microservices MLflow MLOps Prometheus Python SQL Terraform Vertex AI
1 hour, 56 minutes ago

DevOps Support Engineer (Argentina)

Arize AI 51-250 IT Services

Arize AI is hiring an On-Prem Technical Support Engineer in Buenos Aires to support customer-hosted deployments, troubleshoot infrastructure issues, and maintain reliable AI platform environments at scale.

AWS Azure DNS Helm Kubernetes TLS
2 hours, 11 minutes ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers