Xsolla

Xsolla

Xsolla is an international payment solution provider for online games, offering tools to launch, monetize, and scale games worldwide with local payment methods and fraud prevention.

Internet Software & Services
251-1K
Founded 2005

Description

  • Serve as the primary dashboard monitor during shifts and detect anomalies across Datadog signals, logs, metrics, synthetic tests, and RUM.
  • Triage and investigate production incidents, create incident tickets in JIRA Service Management, and route issues to the correct team.
  • Own lower-severity incidents end to end, including diagnosis, runbook execution, resolution, and escalation when needed.
  • Support the TSO Lead during major incidents by surfacing real-time impact data, maintaining incident records, and executing mitigation actions.
  • Draft incident communications, including internal updates, stakeholder notifications, and customer-facing status page messages.
  • Analyze incident trends, recurring issues, and production bugs during non-incident periods and contribute findings to reports.
  • Compile incident timelines, draft initial PIR documents, and track post-incident action items.
  • Build and maintain operational automation, including alert enrichment scripts, incident templates, Slack workflows, and dashboard widgets.
  • Develop and improve runbooks and document repeatable resolution procedures for shift coverage.
  • Conduct structured shift handoffs and participate in knowledge transfer sessions with SREs.
  • Cover for the TSO Lead during absences, including severity classification, escalation decisions, communications, and basic incident commander functions.
  • Publish periodic health reports for critical applications.

Requirements

  • Previous experience working at a gaming company is required.
  • 4+ years of experience in SRE, DevOps, production operations, NOC, or technical operations in a high-availability environment.
  • Strong troubleshooting and investigation skills across logs, APM traces, infrastructure metrics, database queries, and network paths.
  • Hands-on experience with Datadog or an equivalent observability platform such as Grafana, Splunk, New Relic, or Elastic.
  • Proficiency in at least one scripting language: Python, Go, or Bash.
  • Clear written and verbal communication skills in English for incident tickets, updates, handoffs, status pages, and PIR drafts.
  • Working knowledge of Kubernetes and cloud infrastructure; GCP is preferred, with AWS or Azure acceptable.
  • Understanding of SLOs, error budgets, and burn-rate alerting.
  • Experience with incident management tools such as JIRA, JIRA Service Management, PagerDuty, OpsGenie, Slack, and Confluence.
  • Experience with or strong interest in AI/ML-assisted operations such as anomaly detection, alert correlation, predictive monitoring, or automated remediation.
  • Comfort with 24x7 shift-based operations in a follow-the-sun model, including rotating weekend on-call.
  • Familiarity with Datadog Service Catalog, synthetic monitoring, and RUM.
  • Experience debugging distributed systems and tracing failures across microservices.
  • Exposure to database operations involving MySQL, PostgreSQL, Redis, or Kafka.
  • Familiarity with CI/CD and deployment tools such as GitLab CI, ArgoCD, and Helm.
  • JIRA Service Management administration experience, including workflows, automation rules, SLA timers, and queues.
  • ITIL Foundation certification is a plus but not required.

Benefits

  • Unlimited Flexible Time Off.
  • Gym membership.
  • Monthly train ticket.
  • Personalized career roadmap.
  • Training and educational opportunities for professional development.
  • Supportive benefits program focused on physical, mental, and emotional well-being.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Service Reliability Engineer

Thoughtworks 10K-50K Professional Services

Thoughtworks is seeking a Senior Site Reliability Engineer to help clients improve infrastructure reliability, resilience, observability, and operational performance through automation and continuous improvement.

AWS Azure CI/CD Datadog ELK Stack GitOps Grafana Kubernetes Microservices New Relic Nomad Python REST API
3 hours, 20 minutes ago

Cloud Site Reliability Engineer

Cadwell 51-250 Health Care Providers & Services

Cadwell is seeking a Cloud Site Reliability Engineer to operate and improve AWS infrastructure supporting healthcare customers and ensure reliable, secure, and compliant hosted neurodiagnostic software environments.

AWS Bash CI/CD Encryption HIPAA JavaScript JSON Python Terraform TypeScript YAML
1 day, 2 hours ago

Site Reliability Engineer

GiveCampus 51-250 Internet Software & Services

GiveCampus is seeking a hands-on Site Reliability Engineer to strengthen the reliability, performance, observability, and operational maturity of its AWS-based fundraising platform in a remote-first U.S. role.

AWS CI/CD CircleCI Datadog GitHub Actions Kubernetes Linux New Relic OpenSearch PostgreSQL Redis Ruby Ruby on Rails Terraform
1 day, 3 hours ago

DevOps / SRE / DevSecOps Engineer (AWS) - Latin America, Remote

Bluelight Consulting 11-50 Internet Software & Services

Bluelight is hiring a DevOps/SRE/DevSecOps professional to operate secure, reliable AWS-based systems and delivery infrastructure for client software projects in a remote consultancy environment.

AWS CI/CD DevSecOps Docker K6 OpenTelemetry PostgreSQL Secrets Management Terraform
2 days, 2 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers