RUNWARE

RUNWARE

RUNWARE provides an affordable API that enables AI developers to efficiently run image, video, and custom generative AI models without the need for extensive infrastructure or machine learning expertise.

Internet Software & Services
1-10
Founded 2023

Description

  • Own and improve the reliability, availability, and performance of critical production services across the Runware platform.
  • Define and evolve reliability practices, including SLIs, SLOs, alerting, observability, and production-readiness standards.
  • Investigate complex production issues across distributed systems, APIs, networking, queues, databases, and GPU-backed workloads.
  • Participate in the engineering on-call rotation and take issues from initial investigation through long-term remediation.
  • Lead and contribute to incident reviews and root cause analyses, turning recurring failures into engineering improvements.
  • Reduce operational toil through automation, automated remediation, and safer deployment and recovery processes.
  • Work with Engineering and DevOps teams on capacity planning, performance, scaling, and architectural improvements.

Requirements

  • Strong experience operating and troubleshooting production systems at scale in an SRE, Production Engineering, Platform Engineering, or similar role.
  • Strong understanding of distributed systems and comfort debugging across applications, databases, queues, containers, networking, and infrastructure.
  • Experience designing and operating observability systems using metrics, logs, and distributed tracing.
  • Understanding of SRE principles including SLIs, SLOs, error budgets, capacity planning, incident management, and reducing operational toil.
  • Experience with Kubernetes, containers, IaC, and automated deployment practices.
  • Ability to write software and automation using Python, Go, or PHP.
  • Strong ownership of production problems and comfort participating in an engineering on-call rotation.
  • Experience operating high-throughput or low-latency APIs and distributed systems (bonus).
  • Experience with bare-metal infrastructure, GPU environments, or AI and ML workloads (bonus).
  • Experience with RabbitMQ or other distributed messaging and queueing systems (bonus).
  • Experience operating MySQL, Redis, ClickHouse, or similar production data systems (bonus).
  • Experience with global traffic management, load balancing, CDN platforms, and hybrid infrastructure environments (bonus).
  • Experience building automated scaling, capacity management, or self-healing systems (bonus).

Benefits

  • Remote-first work environment with the option to work from home anywhere the company can employ you.
  • Flexible hours outside core collaboration blocks.
  • Generous paid time off, including vacation, sick days, and public holidays.
  • Meaningful stock options.
  • Paid family leave, including maternity, paternity, and caregiver time.
  • Company retreats twice a year.
  • Built-in downtime after big release pushes to rest and recharge.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Site Reliability Engineer III

onXmaps 251-1K Food Products

onX is hiring a Site Reliability Engineer to manage the infrastructure, deployment automation, and observability that help developers ship reliably at scale for its outdoor technology products.

Apache Airflow CockroachDB GCP Kubernetes OpenTelemetry Prometheus SQL Terraform
12 hours, 31 minutes ago

Site Reliability Engineer - South Korea

MinIO 51-250 Internet Software & Services

MinIO is hiring a Site Reliability Engineer to help enhance and operate its cloud-native storage platform for high-performance, scalable, and durable data storage and retrieval.

C C++ GitOps Go Kubernetes Microservices Rust
12 hours, 31 minutes ago

Senior Site Reliability Engineer- FedRamp

Veeam Software 1K-5K Internet Software & Services

Veeam is hiring a Site Reliability Engineer to help build its global SRE function for the Veeam Data Cloud, focused on the Government and Sovereign Cloud environment.

Argo CD Azure Bitbucket C# CI/CD ELK Stack Git GitHub Actions GitLab CI GitOps Go Grafana HIPAA Java JavaScript Kubernetes OpenTelemetry Prometheus Pulumi Terraform TypeScript
13 hours, 31 minutes ago

Senior Site Reliability Engineer (SRE/DevOps)

qode Internet Software & Services

Senior Site Reliability Engineer role at a growing technology company in Vietnam, focused on building reliable, secure, scalable infrastructure and bringing AI systems into production.

Argo CD AWS Azure CI/CD CloudFormation Datadog Flux GCP GitOps Grafana Kubernetes NestJS Node.js Prometheus Pulumi Python Secrets Management Terraform
14 hours, 1 minute ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers