Symphony Solutions

Symphony Solutions

Symphony Solutions is a Cloud and Agile transformation company headquartered in the Netherlands, with delivery centers in Ukraine, Poland, Macedonia, and The Netherlands. Founded in 2008, the company is now 600 people with over 35 international clients...

Internet Software & Services
251-1K
Founded 2008

Description

  • Own production reliability practices together with DevOps and engineering teams.
  • Monitor production health and improve observability across services and infrastructure.
  • Support releases, hotfixes, rollback execution, and post-release watch.
  • Validate deployment readiness and rollback readiness before production changes.
  • Participate in incident response, service recovery, and RCA/postmortem activities.
  • Maintain runbooks, dashboards, alerting rules, and operational documentation.
  • Ensure release evidence is traceable and audit-ready for production changes.
  • Coordinate with Dev, QA, Product, Support, Release Manager, and DevOps during production events.

Requirements

  • Strong hands-on Linux troubleshooting experience with terminal, logs, processes, networking basics, and resource usage.
  • Strong hands-on experience with Kubernetes/GKE, including workloads, pods, services, ingress/gateway, probes, RBAC, resources, autoscaling, and troubleshooting.
  • Practical GCP experience, especially with GKE, IAM, networking, load balancing, Artifact Registry, and production diagnostics.
  • Experience with Docker and containers, including images, registries, runtime debugging, and container lifecycle.
  • Experience with Helm releases, values, deployment state, and rollback.
  • Ability to understand and troubleshoot GitOps workflows using FluxCD, including drift, image automation, and Git-based rollback.
  • CI/CD understanding with the ability to investigate failed pipelines and deployment issues.
  • Experience with observability tools such as Prometheus, Alertmanager, and Grafana, including metrics, logs, traces, dashboards, alerting, and monitoring.
  • Solid networking fundamentals, including DNS, load balancers, ingress, gateways, TLS, routing, and firewall/security rules.
  • Ability to define and apply SLI/SLO/SLA reliability targets.
  • Experience with incident response, including production incidents, rollback, service recovery, and RCA/postmortems.
  • Operational understanding of PostgreSQL, Couchbase, Kafka, and Elasticsearch/ELK or similar systems.
  • Basic security and compliance knowledge, including secrets, IAM/RBAC, audit trails, and release/change evidence.
  • Ability to assess release risk, define monitoring needs, verify production health, and prepare or validate rollback plans.
  • Ability to improve runbooks, incident response procedures, dashboards, alerts, and release checks.
  • Strong ownership, calm under pressure, and clear communication during incidents and production events.
  • Structured troubleshooting, collaboration, documentation discipline, blameless mindset, proactivity, and prioritization skills.
  • Nice to have: Terraform/IaC experience, advanced GCP infrastructure design, ArgoCD or other GitOps tools, advanced database or performance tuning, service mesh experience, on-call rotation experience with PagerDuty/Opsgenie or similar, RCA/postmortem facilitation experience, betting/gaming domain experience, strong written English, and basic understanding of AI concepts such as agents, skills, and MCP.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

Multi Media 51-250 Internet Software & Services

Multi Media, LLC, the company behind Chaturbate, is hiring a remote Site Reliability Engineer to improve the resilience, observability, automation, and performance of its cloud-based live streaming infrastructure.

Ansible Argo CD Bash C C# C++ Django Docker Flask Go Helm Java Kubernetes Laravel Linux Python Rust TCP/IP Terraform
1 hour, 23 minutes ago

Incident Commander

PENN Entertainment 10K-50K Hotels, Restaurants & Leisure

PENN Interactive is hiring an Incident Commander to join its site reliability team and lead incident response and service reliability efforts across its online and physical platforms.

Ansible AWS Docker Elasticsearch GCP Helm JIRA Kafka Kubernetes Linux MySQL PostgreSQL Prometheus Python Redis Terraform
1 day ago

Site Reliability Engineer

MyFitnessPal 10K-50K Health Care Providers & Services

MyFitnessPal is hiring a Software Engineer III, Site Reliability to own production reliability and delivery pipeline security for the PEAS team supporting automation, CI/CD, and self-service platforms.

AWS CI/CD Datadog GitHub Actions Go HIPAA Kubernetes Python Secrets Management Terraform TypeScript
1 day, 1 hour ago

Site Reliability Engineer 2 (Azure)

PhonePe 5K-10K Capital Markets

PhonePe Limited is hiring a Site Reliability Engineer to manage and scale core cloud infrastructure for a high-volume digital payments environment in India.

Ansible Azure Bash DNS Docker Go Grafana HAProxy InfluxDB Java Linux MySQL Nginx Prometheus Python RabbitMQ SaltStack Terraform Ubuntu
2 days ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers