AI infrastructure Engineer (SRE) Bangalore

2 months, 3 weeks ago
Senior
DevOps and Infrastructure
Together

Together

Together AI provides the fastest and most cost-efficient tools for building generative AI models, with a dedicated team of experts to support users in training their own models and advancing AI technology.

IT Services
1-10
$20M raised

Description

  • Participate in an on-call PagerDuty rotation to respond to incidents affecting availability.
  • Build and operate infrastructure using Ansible, Terraform, and Kubernetes to support large-scale concurrency.
  • Develop monitoring systems that maintain high service quality for customers.
  • Design and implement operational processes for deployments, upgrades, and related production workflows.
  • Debug production issues across services and layers of the stack.
  • Identify and drive reliability, performance, and availability improvements in product architecture.
  • Plan and support the growth of Together AI’s infrastructure.

Requirements

  • 7+ years of professional SRE or related experience.
  • Experience operating GPU-enabled Kubernetes clusters or AI infrastructure for machine learning training and inference workloads.
  • Bachelor's degree in Computer Science or a related field, or equivalent work experience.
  • Expert knowledge of Ansible, including roles and playbooks, Terraform, and Kubernetes.
  • Proficiency in programming or scripting languages.
  • Direct experience with monitoring and observability practices.
  • Advanced knowledge of cloud services.
  • Ability to thrive in a collaborative environment with different stakeholders and subject matter experts.
  • Ideally based near Bangalore or working remotely from India.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

OMS Application Engineer

Warner Music Group is seeking an OMS Application Engineer to maintain, upgrade, and support the technical systems powering its global commerce ecosystem, with a focus on reliability, scalability, and long-term stability.

AWS CI/CD GitHub Actions Java Microservices
2 hours, 20 minutes ago

[Job-32080] Senior SRE / Cloud Engineer (Pessoa Engenheira de Plataforma), Brazil

CI&T 5K-10K Internet Software & Services

A CI&T busca uma pessoa Engenheira de Plataforma Sênior para projetar e operar a infraestrutura OCI que sustenta agentes de IA, garantindo alta disponibilidade e baixa latência.

Generative AI Kafka Kubernetes Terraform
2 hours, 20 minutes ago

[Job-32047] Mid-Level AI Engineer

CI&T 5K-10K Internet Software & Services

CI&T is seeking a Mid-Level AI Engineer to help a delivery team build and productionize Generative AI and Agentic AI solutions for a global enterprise client.

CI/CD Docker Generative AI Git Microservices OpenTelemetry Python
2 hours, 35 minutes ago

Sr. Sustaining and Forward Deployed Engineer

Abacus Insights 51-250 Insurance

Abacus Insights is seeking a Senior Site Reliability Engineer – Forward Deployed to operate and improve its AWS- and Databricks-based healthcare data platform while resolving complex production issues and supporting customers.

Apache Spark AWS CI/CD Databricks Kubernetes Python Snowflake
1 day, 3 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers