Site Reliability Engineer

4 days, 19 hours ago
Full-time
Mid Level
DevOps and Infrastructure
Boson AI

Boson AI

Boson AI develops advanced voice agents that utilize foundation models and continuous learning to facilitate seamless and engaging communication between humans and AI, tailored specifically to various business domains.

Internet Software & Services
Founded 2023

Description

  • Design, operate, and improve reliable infrastructure for AI training and inference workloads.
  • Own and automate operational workflows across networking, compute allocation, storage, GPU/server configuration, or AI platforms.
  • Build monitoring, alerting, runbooks, and incident-response practices to improve operability.
  • Diagnose performance, capacity, and reliability issues across hardware, operating systems, networks, schedulers, and distributed workloads.
  • Partner with ML, research, and platform teams to turn workload needs into infrastructure improvements.
  • Improve provisioning, configuration management, testing, and deployment automation.
  • Plan cluster growth, capacity allocation, upgrades, and lifecycle management.
  • Contribute to reliability standards through documentation and post-incident learning.

Requirements

  • 4+ years of experience in site reliability engineering, infrastructure engineering, systems engineering, or a related production-operations role.
  • Strong hands-on expertise in at least one of the following: networking, cluster and systems allocation, distributed storage, GPU and server administration, or AI training/model-serving infrastructure.
  • Experience with networking topics such as firewalls, switching, routing, ASN/BGP configuration, or InfiniBand.
  • Experience with Kubernetes, SLURM, MAAS, or similar cluster management platforms.
  • Experience with distributed storage, particularly Ceph.
  • Experience with GPU and server administration, including CUDA drivers, firmware, BIOS, and hardware troubleshooting.
  • Experience operating production systems with a focus on availability, performance, security, and automation.
  • Strong Linux administration and scripting skills.
  • A systematic approach to troubleshooting across multiple layers of a complex system.
  • Clear written and verbal communication skills and the ability to work effectively with a distributed team.
  • Experience supporting GPU-intensive AI or HPC environments (preferred).
  • Experience with NVIDIA GPUs, CUDA, NCCL, InfiniBand, RDMA, RoCE, or 100Gb+ Ethernet (preferred).
  • Familiarity with Kubernetes, SLURM, MAAS, Terraform, Ansible, or similar infrastructure tooling (preferred).
  • Experience operating or tuning Ceph clusters (preferred).
  • Familiarity with observability tooling such as Prometheus, Grafana, and centralized logging systems (preferred).
  • Experience with hardware provisioning, firmware management, and bare-metal automation (preferred).
  • Experience running large-scale distributed training or high-throughput inference workloads (preferred).
  • Familiarity with cloud and hybrid infrastructure across AWS, GCP, or Azure (preferred).

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

Counterpart Health 51-200 hospital & health care

Counterpart Health is hiring a Senior Site Reliability and Infrastructure Engineer to support and evolve the technology platform behind its primary care tool and maintain reliable infrastructure for domestic and international workloads.

AWS Azure CI/CD Containerd DNS Docker GCP Go gRPC Helm Kubernetes Linux Load Balancing Prometheus Python Shell Scripting TCP/IP
1 day, 19 hours ago

Senior Test Platform & Reliability Engineer - Star Trek Fleet Command

Scopely 1K-5K Internet Software & Services

Scopely is hiring a Senior Test Platform & Reliability Engineer in Ireland to build validation, reliability, and developer enablement platforms for Star Trek Fleet Command’s large-scale live-service backend systems.

AWS Bash CI/CD Docker GitLab Go Python Terraform
1 day, 19 hours ago

Senior Software Engineer - Databases, SRE | Canada | Remote

Grafana 1K-5K IT Services

Grafana Labs is hiring a Senior Software Engineer for its remote SRE team to improve reliability and operability of Grafana Cloud database services for high-SLA customers across AWS, GCP, and Azure.

AWS Azure GCP Go Helm Java Kubernetes Linux Microservices Python Terraform
2 days, 18 hours ago

Senior Site Reliability Engineer

Semios 51-250 Food Products

Semios Group is hiring a Senior Site Reliability Engineer to help scale, secure, and improve the reliability of its global agricultural technology platform.

AWS Azure Bash Buildkite CI/CD Datadog Docker Envoy GCP Git GitHub GitHub Actions GitLab Go Jenkins Kubernetes Linux NATS New Relic Prometheus Python Ruby Splunk Terraform
2 days, 19 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers