RUNWARE

RUNWARE

RUNWARE provides an affordable API that enables AI developers to efficiently run image, video, and custom generative AI models without the need for extensive infrastructure or machine learning expertise.

Internet Software & Services
1-10
Founded 2023

Description

  • Build and scale infrastructure for real-time AI inference across GPU fleets, bare-metal servers, and containerized production systems.
  • Help evolve the platform toward more elastic, on-demand infrastructure that can respond to customer traffic and model demand.
  • Improve the performance, reliability, and resilience of request entrypoints, inference services, queues, storage, load balancers, and networking.
  • Automate infrastructure operations, including provisioning, configuration, CI/CD, deployment safety, progressive rollouts, and rapid rollback.
  • Build and maintain the observability stack needed to detect issues early, understand capacity, and resolve problems before they affect customers.
  • Lead production operations, incident response, debugging, and post-incident improvements.
  • Strengthen infrastructure security and compliance through patching, secrets management, access controls, hardening, auditability, and documentation.

Requirements

  • Strong experience as a DevOps Engineer, SRE, Infrastructure Engineer, Platform Engineer, or in a similar role running production systems at scale.
  • Deep Linux knowledge and confidence debugging real production issues across networking, storage, performance, services, and system behavior.
  • Hands-on experience building automation, Infrastructure-as-Code, CI/CD pipelines, and deployment workflows.
  • Experience operating high-availability, low-latency, or high-throughput platforms where reliability and performance directly affect customers.
  • Strong networking fundamentals across TCP/IP, DNS, load balancing, routing, firewalls, proxies, TLS, and HTTP.
  • A calm and pragmatic approach under pressure, with strong communication, good judgment, and a bias toward automation over manual toil.
  • Experience operating GPU infrastructure for AI/ML inference, including NVIDIA drivers, CUDA, container runtimes, GPU monitoring, capacity planning, and workload isolation (bonus).
  • Familiarity with inference serving and optimization frameworks such as vLLM, TensorRT, Triton, or similar (bonus).

Benefits

  • Remote-first work environment with the option to work from home anywhere they can employ you.
  • Flexible hours outside core collaboration blocks.
  • Generous paid time off, including vacation, sick days, and public holidays.
  • Meaningful stock options.
  • Paid family leave, including maternity, paternity, and caregiver time.
  • Twice-yearly company retreats in inspiring locations.
  • Built-in downtime after major release cycles to unplug and recharge.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

DevOps Engineer - SRE Observability

Lingaro 5K-10K IT Services

An infrastructure-focused role at Lingaro responsible for monitoring, automating, and designing cloud systems within an Azure-based environment.

Azure Azure Pipelines CI/CD Docker GitHub GitHub Actions Grafana Kubernetes MySQL PostgreSQL Prometheus SQL Terraform
21 hours, 37 minutes ago

PnP Pipeline On-call Support Engineer

PHIZENIX 11-50 information technology & services

This role at an internal engineering organization focuses on monitoring and triaging PnP (Power and Performance) pipelines to keep post-silicon work stable and moving efficiently.

CI/CD Jest
1 day, 22 hours ago

Senior Full Stack Engineer

Virtru 51-250 IT Services

Virtru is hiring a Senior Full Stack Engineer to build and operate its digital privacy products and core data protection platform for organizations that manage sensitive data.

AWS Docker GCP Go JavaScript Node.js PagerDuty React Selenium
3 days, 21 hours ago

Staff Forward Deployed Engineer

Tenstorrent 251-1K Internet Software & Services

Tenstorrent is hiring a remote Forward Deployed Engineer in North America to work directly with customers and internal teams on production AI inference deployments for its AI computers.

Grafana Helm Kubernetes OpenTelemetry Prometheus
3 days, 21 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers