Xsolla

Xsolla

Xsolla is an international payment solution provider for online games, offering tools to launch, monetize, and scale games worldwide with local payment methods and fraud prevention.

Internet Software & Services
251-1K
Founded 2005

Description

  • Serve as the primary dashboard monitor during shifts and detect anomalies across Datadog signals, logs, metrics, synthetic tests, and RUM.
  • Triage and investigate production incidents, create incident tickets in JIRA Service Management, and route issues to the correct team.
  • Own lower-severity incidents end to end, including diagnosis, runbook execution, resolution, and escalation when needed.
  • Support the TSO Lead during major incidents by surfacing real-time impact data, maintaining incident records, and executing mitigation actions.
  • Draft incident communications, including internal updates, stakeholder notifications, and customer-facing status page messages.
  • Analyze incident trends, recurring issues, and production bugs during non-incident periods and contribute findings to reports.
  • Compile incident timelines, draft initial PIR documents, and track post-incident action items.
  • Build and maintain operational automation, including alert enrichment scripts, incident templates, Slack workflows, and dashboard widgets.
  • Develop and improve runbooks and document repeatable resolution procedures for shift coverage.
  • Conduct structured shift handoffs and participate in knowledge transfer sessions with SREs.
  • Cover for the TSO Lead during absences, including severity classification, escalation decisions, communications, and basic incident commander functions.
  • Publish periodic health reports for critical applications.

Requirements

  • Previous experience working at a gaming company is required.
  • 4+ years of experience in SRE, DevOps, production operations, NOC, or technical operations in a high-availability environment.
  • Strong troubleshooting and investigation skills across logs, APM traces, infrastructure metrics, database queries, and network paths.
  • Hands-on experience with Datadog or an equivalent observability platform such as Grafana, Splunk, New Relic, or Elastic.
  • Proficiency in at least one scripting language: Python, Go, or Bash.
  • Clear written and verbal communication skills in English for incident tickets, updates, handoffs, status pages, and PIR drafts.
  • Working knowledge of Kubernetes and cloud infrastructure; GCP is preferred, with AWS or Azure acceptable.
  • Understanding of SLOs, error budgets, and burn-rate alerting.
  • Experience with incident management tools such as JIRA, JIRA Service Management, PagerDuty, OpsGenie, Slack, and Confluence.
  • Experience with or strong interest in AI/ML-assisted operations such as anomaly detection, alert correlation, predictive monitoring, or automated remediation.
  • Comfort with 24x7 shift-based operations in a follow-the-sun model, including rotating weekend on-call.
  • Familiarity with Datadog Service Catalog, synthetic monitoring, and RUM.
  • Experience debugging distributed systems and tracing failures across microservices.
  • Exposure to database operations involving MySQL, PostgreSQL, Redis, or Kafka.
  • Familiarity with CI/CD and deployment tools such as GitLab CI, ArgoCD, and Helm.
  • JIRA Service Management administration experience, including workflows, automation rules, SLA timers, and queues.
  • ITIL Foundation certification is a plus but not required.

Benefits

  • Unlimited Flexible Time Off.
  • Gym membership.
  • Monthly train ticket.
  • Personalized career roadmap.
  • Training and educational opportunities for professional development.
  • Supportive benefits program focused on physical, mental, and emotional well-being.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

GoDaddy 5001-10000 Technology, Information and Internet

GoDaddy is hiring a Senior Site Reliability Engineer to support the Commerce ecosystem remotely by improving the reliability, scalability, security, and operation of business-critical production platforms.

Ansible AWS AWS CDK CI/CD CloudFormation Go Kubernetes Linux Pulumi Python SaltStack Terraform TypeScript
3 hours, 56 minutes ago

Senior Site Reliability Engineer

PENN Entertainment 10K-50K Hotels, Restaurants & Leisure

PENN Entertainment is seeking a Senior Site Reliability Engineer to build and operate the cloud infrastructure supporting large-scale sports betting and media platforms across regulated jurisdictions.

Argo CD AWS Bash CI/CD Datadog GCP GitHub Actions GitOps Go Helm Kubernetes Linux PostgreSQL Python Shell Scripting Terraform
1 day, 3 hours ago

Staff Site Reliability Engineer, Ads

Reddit 1K-5K Internet Software & Services

Reddit is hiring a Staff Site Reliability Engineer to provide technical leadership for reliability, scalability, and operational excellence across its advertising infrastructure and revenue-critical systems.

Apache Spark ClickHouse GCP Go Kafka Kubernetes
1 day, 3 hours ago

Sr Lead Network Reliability Engineer

Coupa Software 1K-5K Internet Software & Services

Coupa is hiring a Sr. Lead Network Development Engineer to scale and operate its global SaaS platform’s cloud networking infrastructure through automation, reliability engineering, and technical leadership.

Ansible AWS Azure Chef DNS Fortinet Go Java Kubernetes Linux Python Ruby TCP/IP Terraform TLS
2 days, 3 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers