PENN Entertainment

PENN Entertainment

PENN Entertainment is a leading provider of integrated entertainment, sports content, and casino gaming experiences across North America, offering a diverse range of entertainment destinations and a premier loyalty program that rewards members with cas...

Hotels, Restaurants & Leisure
10K-50K
Founded 1982

Description

  • Lead real-time incident management across P1, P2, P3, and P4 incidents.
  • Classify, document, investigate, escalate, diagnose, and coordinate recovery for incidents.
  • Drive root cause analysis and maintain related documentation.
  • Work cross-functionally with Incident Commanders, Customer Support, Application, Engineering, SRE, Release Engineering, and Infrastructure teams.
  • Improve service delivery and release processes based on disruption reports.
  • Develop and maintain practices, frameworks, process flows, templates, and process guides.
  • Continuously improve internal methodologies, processes, and tools.
  • Communicate incident status and updates to stakeholders via email, Slack, and Teams.
  • Promote Jira release ticket management and alignment with incident communications and SLAs.
  • Identify requirements and gaps that contribute to downtime or blind spots.
  • Support continuous service improvement initiatives.
  • Perform other duties as required.

Requirements

  • Experience in a similar role or incident management role.
  • Experience with containerization, with Docker and Kubernetes preferred.
  • Understanding of configuration management and infrastructure-as-code tools such as Terraform, Ansible, and Helm.
  • Experience with a programming language.
  • Comfortable working in Linux environments.
  • Experience working with AWS, GCP, and on-premises environments.
  • Ability to work independently and learn quickly with little supervision.
  • Ability to handle multiple projects simultaneously.
  • Willingness to drop everything and take on ad hoc tasks.
  • Strong communication skills and ability to extract needed information through conversation.
  • A degree in computer science, engineering, or similar experience.
  • Nice to have: Postgres, MySQL, Elasticsearch, Kafka, Redis, Terragrunt, Prometheus, Python, and Talos Linux.

Benefits

  • Competitive compensation package with salary range of CAD $90,000–$135,000.
  • Fun, relaxed work environment.
  • Education and conference reimbursements.
  • Parental leave top-up.
  • Opportunities for career progression.
  • Mentoring opportunities.
  • Bonus eligibility for most non-sales positions.
  • Best-in-class employee benefits.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Site Reliability Engineer 2 (Azure)

PhonePe 5K-10K Capital Markets

PhonePe Limited is hiring a Site Reliability Engineer to manage and scale core cloud infrastructure for a high-volume digital payments environment in India.

Ansible Azure Bash DNS Docker Go Grafana HAProxy InfluxDB Java Linux MySQL Nginx Prometheus Python RabbitMQ SaltStack Terraform Ubuntu
8 hours, 13 minutes ago

Sr. Control System Engineer/Site Reliability Engineer (SRE)

QuEra Computing 11-50 Internet Software & Services

QuEra is seeking a Sr. Control System Engineer/Site Reliability Engineer to integrate and maintain the hardware and software systems that support its quantum control stack and keep development and production environments reliable.

Ansible Bash CI/CD Debian DHCP DNS Docker ELK Stack Embedded Systems Git GitLab CI Go Grafana Jenkins Kubernetes Linux Prometheus Python TCP/IP Terraform Ubuntu
8 hours, 43 minutes ago

Site Reliability Engineer

VantageScore 11-50 Banks

Site Reliability Engineer at a growing engineering team, focused on DevSecOps for maintaining the reliability, security, and compliance of cloud infrastructure, APIs, and software supply chains.

Agile AWS AWS CDK Bash CI/CD CloudFormation CodePipeline Datadog DevSecOps Docker EC2 GitHub Actions Grafana HashiCorp Vault Kong Kubernetes Microservices Python REST API Scrum Terraform
1 day, 8 hours ago

Application Site Reliability Engineer (SRE)

CXM Direct 51-250 Capital Markets

Application Site Reliability Engineer at a trading technology company, responsible for keeping .NET/C# Windows-based trading and back-office services highly reliable, observable, and resilient.

AWS Bash C# CI/CD Docker Grafana Kubernetes Microservices .NET OpenTelemetry PowerShell Prometheus Python Terraform Windows Server
1 day, 8 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers