Xsolla

Xsolla

Xsolla is an international payment solution provider for online games, offering tools to launch, monetize, and scale games worldwide with local payment methods and fraud prevention.

Internet Software & Services
251-1K
Founded 2005

Description

  • Serve as the primary dashboard monitor during shifts and detect anomalies across Datadog signals, logs, metrics, synthetic tests, and RUM.
  • Triage and investigate production incidents, create incident tickets in JIRA Service Management, and route issues to the correct team.
  • Own lower-severity incidents end to end, including diagnosis, runbook execution, resolution, and escalation when needed.
  • Support the TSO Lead during major incidents by surfacing real-time impact data, maintaining incident records, and executing mitigation actions.
  • Draft incident communications, including internal updates, stakeholder notifications, and customer-facing status page messages.
  • Analyze incident trends, recurring issues, and production bugs during non-incident periods and contribute findings to reports.
  • Compile incident timelines, draft initial PIR documents, and track post-incident action items.
  • Build and maintain operational automation, including alert enrichment scripts, incident templates, Slack workflows, and dashboard widgets.
  • Develop and improve runbooks and document repeatable resolution procedures for shift coverage.
  • Conduct structured shift handoffs and participate in knowledge transfer sessions with SREs.
  • Cover for the TSO Lead during absences, including severity classification, escalation decisions, communications, and basic incident commander functions.
  • Publish periodic health reports for critical applications.

Requirements

  • Previous experience working at a gaming company is required.
  • 4+ years of experience in SRE, DevOps, production operations, NOC, or technical operations in a high-availability environment.
  • Strong troubleshooting and investigation skills across logs, APM traces, infrastructure metrics, database queries, and network paths.
  • Hands-on experience with Datadog or an equivalent observability platform such as Grafana, Splunk, New Relic, or Elastic.
  • Proficiency in at least one scripting language: Python, Go, or Bash.
  • Clear written and verbal communication skills in English for incident tickets, updates, handoffs, status pages, and PIR drafts.
  • Working knowledge of Kubernetes and cloud infrastructure; GCP is preferred, with AWS or Azure acceptable.
  • Understanding of SLOs, error budgets, and burn-rate alerting.
  • Experience with incident management tools such as JIRA, JIRA Service Management, PagerDuty, OpsGenie, Slack, and Confluence.
  • Experience with or strong interest in AI/ML-assisted operations such as anomaly detection, alert correlation, predictive monitoring, or automated remediation.
  • Comfort with 24x7 shift-based operations in a follow-the-sun model, including rotating weekend on-call.
  • Familiarity with Datadog Service Catalog, synthetic monitoring, and RUM.
  • Experience debugging distributed systems and tracing failures across microservices.
  • Exposure to database operations involving MySQL, PostgreSQL, Redis, or Kafka.
  • Familiarity with CI/CD and deployment tools such as GitLab CI, ArgoCD, and Helm.
  • JIRA Service Management administration experience, including workflows, automation rules, SLA timers, and queues.
  • ITIL Foundation certification is a plus but not required.

Benefits

  • Unlimited Flexible Time Off.
  • Gym membership.
  • Monthly train ticket.
  • Personalized career roadmap.
  • Training and educational opportunities for professional development.
  • Supportive benefits program focused on physical, mental, and emotional well-being.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Software/Site Reliability Engineer - FedRAMP

Tenable 1K-5K Internet Software & Services

Tenable is hiring a Site Reliability Engineer to help scale and operate its cloud-based vulnerability management platform for private and U.S. Government cloud customers.

Agile AWS Azure Bash CI/CD Datadog Docker DynamoDB Elasticsearch GCP Go Gradle Groovy Helm Java Kafka Kotlin Kubernetes Microservices Node.js OpenSearch OpenTelemetry Python Splunk Terraform
9 hours, 9 minutes ago

Senior Site Reliability Engineer (SRE)

Branch 51-250 Professional Services

Branch is hiring a Senior Site Reliability Engineer to improve the reliability, scalability, performance, and observability of its fintech platform through automation and operational best practices.

Bash Docker GCP Go Gradle Grafana Java Kubernetes MySQL OpenTelemetry Prometheus Python Redis Spring Boot Terraform
9 hours, 39 minutes ago

Staff Site Reliability Engineer

Filevine 251-1K Specialized Consumer Services

Filevine is seeking a Staff Site Reliability Engineer to lead reliability strategy and production excellence for its cloud platform and distributed systems supporting modern legal operations.

Bash Datadog Go HIPAA Kubernetes Machine Learning New Relic Python
1 day, 10 hours ago

Senior Site Reliability Engineer – Telephony & Communications Platform (AWS)

Filevine 251-1K Specialized Consumer Services

Filevine is hiring a Site Reliability Engineer to strengthen the reliability, scalability, and recoverability of its legal AI platform and the systems that support it.

AWS Azure Bash CI/CD EC2 GCP HIPAA PowerShell Python Terraform Twilio
1 day, 10 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers