RUNWARE

RUNWARE

RUNWARE provides an affordable API that enables AI developers to efficiently run image, video, and custom generative AI models without the need for extensive infrastructure or machine learning expertise.

Internet Software & Services
1-10
Founded 2023

Description

  • Own and improve the reliability, availability, and performance of critical production services across the Runware platform.
  • Define and evolve reliability practices, including SLIs, SLOs, alerting, observability, and production-readiness standards.
  • Investigate complex production issues across distributed systems, APIs, networking, queues, databases, and GPU-backed workloads.
  • Participate in the engineering on-call rotation and take issues from initial investigation through long-term remediation.
  • Lead and contribute to incident reviews and root cause analyses, turning recurring failures into engineering improvements.
  • Reduce operational toil through automation, automated remediation, and safer deployment and recovery processes.
  • Work with Engineering and DevOps teams on capacity planning, performance, scaling, and architectural improvements.

Requirements

  • Strong experience operating and troubleshooting production systems at scale in an SRE, Production Engineering, Platform Engineering, or similar role.
  • Strong understanding of distributed systems and comfort debugging across applications, databases, queues, containers, networking, and infrastructure.
  • Experience designing and operating observability systems using metrics, logs, and distributed tracing.
  • Understanding of SRE principles including SLIs, SLOs, error budgets, capacity planning, incident management, and reducing operational toil.
  • Experience with Kubernetes, containers, IaC, and automated deployment practices.
  • Ability to write software and automation using Python, Go, or PHP.
  • Strong ownership of production problems and comfort participating in an engineering on-call rotation.
  • Experience operating high-throughput or low-latency APIs and distributed systems (bonus).
  • Experience with bare-metal infrastructure, GPU environments, or AI and ML workloads (bonus).
  • Experience with RabbitMQ or other distributed messaging and queueing systems (bonus).
  • Experience operating MySQL, Redis, ClickHouse, or similar production data systems (bonus).
  • Experience with global traffic management, load balancing, CDN platforms, and hybrid infrastructure environments (bonus).
  • Experience building automated scaling, capacity management, or self-healing systems (bonus).

Benefits

  • Remote-first work environment with the option to work from home anywhere the company can employ you.
  • Flexible hours outside core collaboration blocks.
  • Generous paid time off, including vacation, sick days, and public holidays.
  • Meaningful stock options.
  • Paid family leave, including maternity, paternity, and caregiver time.
  • Company retreats twice a year.
  • Built-in downtime after big release pushes to rest and recharge.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

PandaDoc 251-1K Internet Software & Services

PandaDoc is hiring a Site Reliability Engineer to help keep its document workflow platform reliable, resilient, and available for customers.

Agile AWS Django Grafana Java Kafka Kubernetes NATS PostgreSQL Python RabbitMQ Spring Boot
15 hours, 52 minutes ago

Senior Site Reliability Engineer

PandaDoc 251-1K Internet Software & Services

PandaDoc is hiring a Site Reliability Engineer to help keep its document workflow platform reliable through incident management, observability, production support, and resilience work across services.

Agile AWS Django Grafana Java Kafka Kubernetes NATS PostgreSQL Python RabbitMQ Spring Boot
15 hours, 52 minutes ago

Senior Site Reliability Engineer

PandaDoc 251-1K Internet Software & Services

PandaDoc is hiring a Site Reliability Engineer to keep its document workflow platform highly available and resilient while supporting production operations and reliability improvements.

Agile AWS Django Grafana Java Kafka Kubernetes NATS PostgreSQL Python RabbitMQ Spring Boot
1 day, 15 hours ago

Senior Site Reliability Engineer

PandaDoc 251-1K Internet Software & Services

PandaDoc is hiring a Site Reliability Engineer to help keep its production document workflow platform reliable, resilient, and low-downtime for customers.

AWS Django Grafana Java Kafka Kubernetes NATS PostgreSQL Python RabbitMQ Spring Boot
1 day, 15 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers