RUNWARE

RUNWARE

RUNWARE provides an affordable API that enables AI developers to efficiently run image, video, and custom generative AI models without the need for extensive infrastructure or machine learning expertise.

Internet Software & Services
1-10
Founded 2023

Description

  • Build and scale infrastructure for real-time AI inference across GPU fleets, bare-metal servers, and containerized production systems.
  • Help evolve the platform toward more elastic, on-demand infrastructure that can respond to customer traffic and model demand.
  • Improve the performance, reliability, and resilience of request entrypoints, inference services, queues, storage, load balancers, and networking.
  • Automate infrastructure operations, including provisioning, configuration, CI/CD, deployment safety, progressive rollouts, and rapid rollback.
  • Build and maintain the observability stack needed to detect issues early, understand capacity, and resolve problems before they affect customers.
  • Lead production operations, incident response, debugging, and post-incident improvements.
  • Strengthen infrastructure security and compliance through patching, secrets management, access controls, hardening, auditability, and documentation.

Requirements

  • Strong experience as a DevOps Engineer, SRE, Infrastructure Engineer, Platform Engineer, or in a similar role running production systems at scale.
  • Deep Linux knowledge and confidence debugging real production issues across networking, storage, performance, services, and system behavior.
  • Hands-on experience building automation, Infrastructure-as-Code, CI/CD pipelines, and deployment workflows.
  • Experience operating high-availability, low-latency, or high-throughput platforms where reliability and performance directly affect customers.
  • Strong networking fundamentals across TCP/IP, DNS, load balancing, routing, firewalls, proxies, TLS, and HTTP.
  • A calm and pragmatic approach under pressure, with strong communication, good judgment, and a bias toward automation over manual toil.
  • Experience operating GPU infrastructure for AI/ML inference, including NVIDIA drivers, CUDA, container runtimes, GPU monitoring, capacity planning, and workload isolation (bonus).
  • Familiarity with inference serving and optimization frameworks such as vLLM, TensorRT, Triton, or similar (bonus).

Benefits

  • Remote-first work environment with the option to work from home anywhere they can employ you.
  • Flexible hours outside core collaboration blocks.
  • Generous paid time off, including vacation, sick days, and public holidays.
  • Meaningful stock options.
  • Paid family leave, including maternity, paternity, and caregiver time.
  • Twice-yearly company retreats in inspiring locations.
  • Built-in downtime after major release cycles to unplug and recharge.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

DevOps Engineer

Arize AI 51-250 IT Services

Arize AI is hiring an On-Prem Infrastructure/DevOps Engineer to deploy, operate, and scale its AI observability platform across SaaS and customer-hosted environments.

AWS Azure Kubernetes
14 hours, 37 minutes ago

DevOps Support Engineer (Argentina)

Arize AI 51-250 IT Services

Arize AI is hiring an On-Prem Technical Support Engineer in Buenos Aires to support customer-hosted deployments, troubleshoot infrastructure issues, and maintain reliable AI platform environments at scale.

AWS Azure DNS Helm Kubernetes TLS
14 hours, 52 minutes ago

Senior DevOps Engineer

Softeta 51-250 Internet Software & Services

Softeta is seeking an on-premises DevOps Engineer to help an e-commerce client migrate from public cloud infrastructure, operate high-traffic systems, and improve infrastructure reliability and observability.

Agile Ansible Argo CD CI/CD GCP Go Grafana HAProxy Kafka Kubernetes Linux Lua Prometheus Redis Scrum Terraform Zabbix
15 hours, 37 minutes ago

Developer Advocate - Service Management EMEA

Datadog 5K-10K IT Services

Datadog is seeking a service-management and technical advocacy engineer to help SRE, DevOps, and operations communities improve incident response, observability, and operational automation through engineering and technical storytelling.

Bash Datadog Go Node.js OpsGenie PagerDuty Python
1 day, 14 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers