RUNWARE

RUNWARE

RUNWARE provides an affordable API that enables AI developers to efficiently run image, video, and custom generative AI models without the need for extensive infrastructure or machine learning expertise.

Internet Software & Services
1-10
Founded 2023

Description

  • Own and improve the reliability, availability, and performance of critical production services across the Runware platform.
  • Define and evolve reliability practices, including SLIs, SLOs, alerting, observability, and production-readiness standards.
  • Investigate complex production issues across distributed systems, APIs, networking, queues, databases, and GPU-backed workloads.
  • Participate in the engineering on-call rotation and take issues from initial investigation through long-term remediation.
  • Lead and contribute to incident reviews and root cause analyses, turning recurring failures into engineering improvements.
  • Reduce operational toil through automation, automated remediation, and safer deployment and recovery processes.
  • Work with Engineering and DevOps teams on capacity planning, performance, scaling, and architectural improvements.

Requirements

  • Strong experience operating and troubleshooting production systems at scale in an SRE, Production Engineering, Platform Engineering, or similar role.
  • Strong understanding of distributed systems and comfort debugging across applications, databases, queues, containers, networking, and infrastructure.
  • Experience designing and operating observability systems using metrics, logs, and distributed tracing.
  • Understanding of SRE principles including SLIs, SLOs, error budgets, capacity planning, incident management, and reducing operational toil.
  • Experience with Kubernetes, containers, IaC, and automated deployment practices.
  • Ability to write software and automation using Python, Go, or PHP.
  • Strong ownership of production problems and comfort participating in an engineering on-call rotation.
  • Experience operating high-throughput or low-latency APIs and distributed systems (bonus).
  • Experience with bare-metal infrastructure, GPU environments, or AI and ML workloads (bonus).
  • Experience with RabbitMQ or other distributed messaging and queueing systems (bonus).
  • Experience operating MySQL, Redis, ClickHouse, or similar production data systems (bonus).
  • Experience with global traffic management, load balancing, CDN platforms, and hybrid infrastructure environments (bonus).
  • Experience building automated scaling, capacity management, or self-healing systems (bonus).

Benefits

  • Remote-first work environment with the option to work from home anywhere the company can employ you.
  • Flexible hours outside core collaboration blocks.
  • Generous paid time off, including vacation, sick days, and public holidays.
  • Meaningful stock options.
  • Paid family leave, including maternity, paternity, and caregiver time.
  • Company retreats twice a year.
  • Built-in downtime after big release pushes to rest and recharge.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Service Reliability Engineer

Thoughtworks 10K-50K Professional Services

Thoughtworks is seeking a Senior Site Reliability Engineer to help clients improve infrastructure reliability, resilience, observability, and operational performance through automation and continuous improvement.

AWS Azure CI/CD Datadog ELK Stack GitOps Grafana Kubernetes Microservices New Relic Nomad Python REST API
15 hours, 37 minutes ago

Cloud Site Reliability Engineer

Cadwell 51-250 Health Care Providers & Services

Cadwell is seeking a Cloud Site Reliability Engineer to operate and improve AWS infrastructure supporting healthcare customers and ensure reliable, secure, and compliant hosted neurodiagnostic software environments.

AWS Bash CI/CD Encryption HIPAA JavaScript JSON Python Terraform TypeScript YAML
1 day, 15 hours ago

Site Reliability Engineer

GiveCampus 51-250 Internet Software & Services

GiveCampus is seeking a hands-on Site Reliability Engineer to strengthen the reliability, performance, observability, and operational maturity of its AWS-based fundraising platform in a remote-first U.S. role.

AWS CI/CD CircleCI Datadog GitHub Actions Kubernetes Linux New Relic OpenSearch PostgreSQL Redis Ruby Ruby on Rails Terraform
1 day, 15 hours ago

DevOps / SRE / DevSecOps Engineer (AWS) - Latin America, Remote

Bluelight Consulting 11-50 Internet Software & Services

Bluelight is hiring a DevOps/SRE/DevSecOps professional to operate secure, reliable AWS-based systems and delivery infrastructure for client software projects in a remote consultancy environment.

AWS CI/CD DevSecOps Docker K6 OpenTelemetry PostgreSQL Secrets Management Terraform
2 days, 15 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers