Yuno

Yuno is a payment orchestration and financial infrastructure platform for global businesses. It centralizes payment methods, PSPs, fraud tools, routing, reconciliation, checkout, payouts, and other payment workflows through a single API.

Payment Processing Software
51-200
Founded 2022

Description

  • Define the reliability strategy, including SLO culture, error-budget policy, and incident practices across engineering teams.
  • Drive architectural decisions as the platform evolves and determine when infrastructure changes are needed.
  • Design and own the messaging layer for reliable inter-service communication and async event-driven systems.
  • Own cloud infrastructure and automate provisioning with infrastructure as code.
  • Build monitoring, tracing, and alerting to improve platform observability and incident detection.
  • Serve as the senior escalation point for complex production issues and lead incident response.
  • Run blameless postmortems and root-cause analyses that result in permanent fixes.
  • Mentor senior and mid-level engineers to raise the reliability bar across the organization.
  • Perform chaos engineering and resilience experiments to identify failure modes before they reach production.

Requirements

  • 7+ years of experience.
  • Experience designing and owning event-driven architectures and messaging systems such as Kafka, NATS, or RabbitMQ.
  • Deep AWS experience with EC2, VPC, IAM, S3, and RDS.
  • Strong networking fundamentals.
  • Hands-on experience with infrastructure as code tools such as Terraform or Pulumi.
  • Production experience with Kubernetes and Docker.
  • Experience with observability tools such as Datadog, including dashboards, monitors, APM, and distributed tracing.
  • Track record defining and operating SLOs, SLIs, and error budgets.
  • Hands-on chaos engineering or resilience testing experience using tools such as Gremlin, Chaos Mesh, or AWS FIS.
  • Experience debugging distributed systems and cascading production failures.
  • Comfort coding automation and tooling in Go, Python, or similar languages.
  • Solid SQL experience with PostgreSQL and NoSQL experience with MongoDB and Redis.
  • Proven technical leadership influencing architecture and reliability standards across teams.
  • Advanced English proficiency, written and spoken.
  • Preferred: experience with AI/MLOps infrastructure, including model serving, LLM inference, GPU/resource management, and tools like LangFuse, LangSmith, Braintrust, or MLflow.
  • Preferred: experience with multi-tenant container platforms such as Replit, Railway, Fly.io, or internal PaaS.
  • Preferred: experience with data pipelines and orchestration tools such as Airflow or Prefect, and warehouses such as Databricks, Snowflake, or BigQuery.
  • Preferred: experience with incident management and on-call tools such as PagerDuty, Opsgenie, or incident.io.
  • Preferred: experience in the payments industry.
  • Nice to have: ECS experience.
  • Nice to have: experience with s6-overlay for container process supervision.
  • Nice to have: experience with AI agent framework ecosystems.
  • Nice to have: Spanish proficiency.

Benefits

  • Competitive compensation.
  • Remote work from anywhere.
  • One-time home office bonus.
  • Work equipment provided.
  • Stock options.
  • Health plan wherever you are.
  • Flexible days off.
  • Language, professional, and personal growth courses.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer (SRE)

UJET 251-1K Professional Services

UJET is seeking a Senior Site Reliability Engineer to build and scale its SRE function, improving reliability, reducing operational toil, and establishing production engineering best practices for its cloud-based contact center platform.

AWS Azure Go Java Kubernetes Microservices Python Terraform
3 hours, 15 minutes ago

[31893] AI ENGINEER SRE (APP)

CI&T 5K-10K Internet Software & Services

Engenheiro(a) SRE/Software Sênior na CI&T, atuando na rearquitetura e melhoria de performance de aplicativos de saúde, desde a definição da arquitetura até a entrega de soluções testadas para produção.

Angular AWS CI/CD Confluence Java JIRA Kanban Microservices Node.js Scrum Spring Boot
1 day, 3 hours ago

Graduate Site Reliability Engineer

Axon 1K-5K Professional Services

Axon is seeking a junior Site Reliability Engineer in Australia to improve the reliability, scalability, security, and operational performance of mission-critical cloud-native services.

AWS Azure C# CI/CD Go Java Kubernetes Linux Python Terraform
1 day, 4 hours ago

SRE / Performance Engineer - Argentina

Coderio 51-250 Internet Software & Services

Coderio busca un/a SRE/Performance Engineer en Argentina para un entorno bancario, responsable de configurar, monitorear y analizar soluciones destinadas a pruebas de performance.

AWS Grafana Java Jenkins Kafka Kubernetes Microservices .NET Node.js OpenShift Prometheus Python RabbitMQ
1 day, 4 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers