Yuno

Yuno is a payment orchestration and financial infrastructure platform for global businesses. It centralizes payment methods, PSPs, fraud tools, routing, reconciliation, checkout, payouts, and other payment workflows through a single API.

Payment Processing Software
51-200
Founded 2022

Description

  • Define the reliability strategy, including SLO culture, error-budget policy, and incident practices across engineering teams.
  • Drive architectural decisions as the platform evolves and determine when infrastructure changes are needed.
  • Design and own the messaging layer for reliable inter-service communication and async event-driven systems.
  • Own cloud infrastructure and automate provisioning with infrastructure as code.
  • Build monitoring, tracing, and alerting to improve platform observability and incident detection.
  • Serve as the senior escalation point for complex production issues and lead incident response.
  • Run blameless postmortems and root-cause analyses that result in permanent fixes.
  • Mentor senior and mid-level engineers to raise the reliability bar across the organization.
  • Perform chaos engineering and resilience experiments to identify failure modes before they reach production.

Requirements

  • 7+ years of experience.
  • Experience designing and owning event-driven architectures and messaging systems such as Kafka, NATS, or RabbitMQ.
  • Deep AWS experience with EC2, VPC, IAM, S3, and RDS.
  • Strong networking fundamentals.
  • Hands-on experience with infrastructure as code tools such as Terraform or Pulumi.
  • Production experience with Kubernetes and Docker.
  • Experience with observability tools such as Datadog, including dashboards, monitors, APM, and distributed tracing.
  • Track record defining and operating SLOs, SLIs, and error budgets.
  • Hands-on chaos engineering or resilience testing experience using tools such as Gremlin, Chaos Mesh, or AWS FIS.
  • Experience debugging distributed systems and cascading production failures.
  • Comfort coding automation and tooling in Go, Python, or similar languages.
  • Solid SQL experience with PostgreSQL and NoSQL experience with MongoDB and Redis.
  • Proven technical leadership influencing architecture and reliability standards across teams.
  • Advanced English proficiency, written and spoken.
  • Preferred: experience with AI/MLOps infrastructure, including model serving, LLM inference, GPU/resource management, and tools like LangFuse, LangSmith, Braintrust, or MLflow.
  • Preferred: experience with multi-tenant container platforms such as Replit, Railway, Fly.io, or internal PaaS.
  • Preferred: experience with data pipelines and orchestration tools such as Airflow or Prefect, and warehouses such as Databricks, Snowflake, or BigQuery.
  • Preferred: experience with incident management and on-call tools such as PagerDuty, Opsgenie, or incident.io.
  • Preferred: experience in the payments industry.
  • Nice to have: ECS experience.
  • Nice to have: experience with s6-overlay for container process supervision.
  • Nice to have: experience with AI agent framework ecosystems.
  • Nice to have: Spanish proficiency.

Benefits

  • Competitive compensation.
  • Remote work from anywhere.
  • One-time home office bonus.
  • Work equipment provided.
  • Stock options.
  • Health plan wherever you are.
  • Flexible days off.
  • Language, professional, and personal growth courses.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Vice President, Global Production Operations & Reliability

Everbridge 1K-5K Internet Software & Services

Everbridge is hiring a Vice President, Global Production Operations & Reliability to lead the company’s global production operations for its cloud-native SaaS platform and drive reliability, scalability, security, and operational excellence.

AWS CI/CD Kubernetes
20 hours, 6 minutes ago

DevOps Engineer - SRE Observability

Lingaro 5K-10K IT Services

An infrastructure-focused role at Lingaro responsible for monitoring, automating, and designing cloud systems within an Azure-based environment.

Azure Azure Pipelines CI/CD Docker GitHub GitHub Actions Grafana Kubernetes MySQL PostgreSQL Prometheus SQL Terraform
20 hours, 6 minutes ago

Staff Field Reliability Engineer

Honeycomb.io 51-250 Internet Software & Services

Honeycomb is hiring a Field Reliability Engineer to lead complex customer escalations, managed infrastructure operations, and observability strategy for its cloud-based platform.

Ansible AWS Chef EC2 Go Helm Honeycomb Java Kubernetes .NET Node.js OpenTelemetry Python Serverless Terraform TypeScript
20 hours, 21 minutes ago

Senior Site Reliability Engineer

PandaDoc 251-1K Internet Software & Services

PandaDoc is hiring a Site Reliability Engineer to help keep its document workflow platform reliable, resilient, and available for customers.

Agile AWS Django Grafana Java Kafka Kubernetes NATS PostgreSQL Python RabbitMQ Spring Boot
1 day, 20 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers