RUNWARE

RUNWARE

RUNWARE provides an affordable API that enables AI developers to efficiently run image, video, and custom generative AI models without the need for extensive infrastructure or machine learning expertise.

Internet Software & Services
1-10
Founded 2023

Description

  • Own and improve the reliability, availability, and performance of critical production services across the Runware platform.
  • Define and evolve reliability practices, including SLIs, SLOs, alerting, observability, and production-readiness standards.
  • Investigate complex production issues across distributed systems, APIs, networking, queues, databases, and GPU-backed workloads.
  • Participate in the engineering on-call rotation and take issues from initial investigation through long-term remediation.
  • Lead and contribute to incident reviews and root cause analyses, turning recurring failures into engineering improvements.
  • Reduce operational toil through automation, automated remediation, and safer deployment and recovery processes.
  • Work with Engineering and DevOps teams on capacity planning, performance, scaling, and architectural improvements.

Requirements

  • Strong experience operating and troubleshooting production systems at scale in an SRE, Production Engineering, Platform Engineering, or similar role.
  • Strong understanding of distributed systems and comfort debugging across applications, databases, queues, containers, networking, and infrastructure.
  • Experience designing and operating observability systems using metrics, logs, and distributed tracing.
  • Understanding of SRE principles including SLIs, SLOs, error budgets, capacity planning, incident management, and reducing operational toil.
  • Experience with Kubernetes, containers, IaC, and automated deployment practices.
  • Ability to write software and automation using Python, Go, or PHP.
  • Strong ownership of production problems and comfort participating in an engineering on-call rotation.
  • Experience operating high-throughput or low-latency APIs and distributed systems (bonus).
  • Experience with bare-metal infrastructure, GPU environments, or AI and ML workloads (bonus).
  • Experience with RabbitMQ or other distributed messaging and queueing systems (bonus).
  • Experience operating MySQL, Redis, ClickHouse, or similar production data systems (bonus).
  • Experience with global traffic management, load balancing, CDN platforms, and hybrid infrastructure environments (bonus).
  • Experience building automated scaling, capacity management, or self-healing systems (bonus).

Benefits

  • Remote-first work environment with the option to work from home anywhere the company can employ you.
  • Flexible hours outside core collaboration blocks.
  • Generous paid time off, including vacation, sick days, and public holidays.
  • Meaningful stock options.
  • Paid family leave, including maternity, paternity, and caregiver time.
  • Company retreats twice a year.
  • Built-in downtime after big release pushes to rest and recharge.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Specialist II

Everbridge 1K-5K Internet Software & Services

Everbridge is hiring a Senior Site Reliability Engineer II to build resilient platforms and improve the reliability, scalability, and operational readiness of systems supporting critical-event communications.

CI/CD Kubernetes Linux
10 hours, 4 minutes ago

Site Reliability Engineer (SRE)

Rocket.net 11-50 IT Services

Rocket.net is seeking a Site Reliability Engineer to maintain the reliability and performance of its hosting platform while resolving complex infrastructure issues and providing advanced support to customers.

Apache Bash CDN Cloudflare Datadog DNS Linux MariaDB MySQL Nginx Redis SSH WAF WordPress
10 hours, 4 minutes ago

Senior Site Reliability Engineer

GoDaddy 5001-10000 Technology, Information and Internet

GoDaddy is hiring a Senior Site Reliability Engineer to support the Commerce ecosystem remotely by improving the reliability, scalability, security, and operation of business-critical production platforms.

Ansible AWS AWS CDK CI/CD CloudFormation Go Kubernetes Linux Pulumi Python SaltStack Terraform TypeScript
1 day, 9 hours ago

Senior Site Reliability Engineer

PENN Entertainment 10K-50K Hotels, Restaurants & Leisure

PENN Entertainment is seeking a Senior Site Reliability Engineer to build and operate the cloud infrastructure supporting large-scale sports betting and media platforms across regulated jurisdictions.

Argo CD AWS Bash CI/CD Datadog GCP GitHub Actions GitOps Go Helm Kubernetes Linux PostgreSQL Python Shell Scripting Terraform
2 days, 9 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers