Xsolla

Xsolla

Xsolla is an international payment solution provider for online games, offering tools to launch, monetize, and scale games worldwide with local payment methods and fraud prevention.

Internet Software & Services
251-1K
Founded 2005

Description

  • Own the Monetization domain’s application-level infrastructure, including Helm charts, Terraform configurations, Kubernetes deployments, runtime configuration, and service networking/integrations.
  • Design and implement observability for critical services, including SLOs/SLIs, monitors, alerts, and dashboards using Datadog and OpenTelemetry-based tooling.
  • Help set up and improve CI/CD pipelines, including deploy and rollback automation.
  • Perform capacity planning and performance tuning for launches, sales events, and regional rollouts, including load testing and regression investigation.
  • Run Production Readiness Reviews and define what production-ready means for the domain.
  • Support incident response, including deep incident investigation, post-mortems, follow-up reliability improvements, and runbook maintenance.
  • Build automation to reduce operational toil, such as runbook automation, deploy helpers, and operational scripts.
  • Maintain a forward-looking reliability roadmap with product engineering leads.
  • Participate in product planning, refinements, and architecture reviews to bring reliability considerations in early.
  • Co-author company-wide SLO/SLI, capacity, and operational standards with the broader SRE team and contribute to shared SRE subsystems.
  • Participate in the SRE duty rotation and support developers across the company.

Requirements

  • 3+ years of proven SRE, DevOps, or platform engineering experience with on-call/incident response, SLO/monitoring ownership, and deploy/infrastructure work for production services.
  • Software development background with experience building and shipping backend services, not just operating them.
  • Comfort reading application code during investigations and writing production-quality automation in at least one language such as Go or PHP.
  • Hands-on Kubernetes experience, including Helm, manifests, deploy strategies, and debugging application-level performance and networking issues.
  • Experience with managed Kubernetes such as GKE, or comparable platforms.
  • Solid observability experience building monitors, dashboards, and SLOs/SLIs on tools such as Datadog, with Prometheus/Grafana also relevant.
  • Familiarity with OpenTelemetry.
  • Infrastructure as Code experience with Terraform or Terragrunt.
  • GCP experience, including IAM, networking, and managed services.
  • Experience building and maintaining CI/CD pipelines with GitLab CI and/or GitHub Actions.
  • Programming/scripting proficiency in tools such as Python, Go, or Bash.
  • Practical experience with incident response, post-mortems, and driving reliability improvements.
  • Strong collaboration and communication skills for working embedded with product teams.
  • Experience in payments, fintech, e-commerce, or gaming/high-traffic transactional systems is preferred.
  • Kubernetes certifications, Google Cloud Platform certifications, or HashiCorp certifications are nice to have.

Benefits

  • Medical, dental, and vision coverage.
  • PTO.
  • A personalized career roadmap for each employee.
  • Training and educational opportunities for professional development.
  • A supportive benefits program focused on employees’ physical, mental, and emotional well-being.
  • Equal opportunity employment in an inclusive workplace.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

Counterpart Health 51-200 hospital & health care

Counterpart Health is hiring a Senior Site Reliability and Infrastructure Engineer to support and evolve the technology platform behind its primary care tool and maintain reliable infrastructure for domestic and international workloads.

AWS Azure CI/CD Containerd DNS Docker GCP Go gRPC Helm Kubernetes Linux Load Balancing Prometheus Python Shell Scripting TCP/IP
15 hours, 1 minute ago

Senior Test Platform & Reliability Engineer - Star Trek Fleet Command

Scopely 1K-5K Internet Software & Services

Scopely is hiring a Senior Test Platform & Reliability Engineer in Ireland to build validation, reliability, and developer enablement platforms for Star Trek Fleet Command’s large-scale live-service backend systems.

AWS Bash CI/CD Docker GitLab Go Python Terraform
15 hours, 16 minutes ago

Senior Software Engineer - Databases, SRE | Canada | Remote

Grafana 1K-5K IT Services

Grafana Labs is hiring a Senior Software Engineer for its remote SRE team to improve reliability and operability of Grafana Cloud database services for high-SLA customers across AWS, GCP, and Azure.

AWS Azure GCP Go Helm Java Kubernetes Linux Microservices Python Terraform
1 day, 14 hours ago

Senior Site Reliability Engineer

Semios 51-250 Food Products

Semios Group is hiring a Senior Site Reliability Engineer to help scale, secure, and improve the reliability of its global agricultural technology platform.

AWS Azure Bash Buildkite CI/CD Datadog Docker Envoy GCP Git GitHub GitHub Actions GitLab Go Jenkins Kubernetes Linux NATS New Relic Prometheus Python Ruby Splunk Terraform
1 day, 15 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers