PENN Entertainment

PENN Entertainment

PENN Entertainment is a leading provider of integrated entertainment, sports content, and casino gaming experiences across North America, offering a diverse range of entertainment destinations and a premier loyalty program that rewards members with cas...

Hotels, Restaurants & Leisure
10K-50K
Founded 1982

Description

  • Lead real-time incident management across P1, P2, P3, and P4 incidents.
  • Classify, document, investigate, escalate, diagnose, and coordinate recovery for incidents.
  • Drive root cause analysis and maintain related documentation.
  • Work cross-functionally with Incident Commanders, Customer Support, Application, Engineering, SRE, Release Engineering, and Infrastructure teams.
  • Improve service delivery and release processes based on disruption reports.
  • Develop and maintain practices, frameworks, process flows, templates, and process guides.
  • Continuously improve internal methodologies, processes, and tools.
  • Communicate incident status and updates to stakeholders via email, Slack, and Teams.
  • Promote Jira release ticket management and alignment with incident communications and SLAs.
  • Identify requirements and gaps that contribute to downtime or blind spots.
  • Support continuous service improvement initiatives.
  • Perform other duties as required.

Requirements

  • Experience in a similar role or incident management role.
  • Experience with containerization, with Docker and Kubernetes preferred.
  • Understanding of configuration management and infrastructure-as-code tools such as Terraform, Ansible, and Helm.
  • Experience with a programming language.
  • Comfortable working in Linux environments.
  • Experience working with AWS, GCP, and on-premises environments.
  • Ability to work independently and learn quickly with little supervision.
  • Ability to handle multiple projects simultaneously.
  • Willingness to drop everything and take on ad hoc tasks.
  • Strong communication skills and ability to extract needed information through conversation.
  • A degree in computer science, engineering, or similar experience.
  • Nice to have: Postgres, MySQL, Elasticsearch, Kafka, Redis, Terragrunt, Prometheus, Python, and Talos Linux.

Benefits

  • Competitive compensation package with salary range of CAD $90,000–$135,000.
  • Fun, relaxed work environment.
  • Education and conference reimbursements.
  • Parental leave top-up.
  • Opportunities for career progression.
  • Mentoring opportunities.
  • Bonus eligibility for most non-sales positions.
  • Best-in-class employee benefits.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

GoDaddy 5001-10000 Technology, Information and Internet

GoDaddy is hiring a Senior Site Reliability Engineer to support the Commerce ecosystem remotely by improving the reliability, scalability, security, and operation of business-critical production platforms.

Ansible AWS AWS CDK CI/CD CloudFormation Go Kubernetes Linux Pulumi Python SaltStack Terraform TypeScript
57 minutes ago

Staff Site Reliability Engineer, Ads

Reddit 1K-5K Internet Software & Services

Reddit is hiring a Staff Site Reliability Engineer to provide technical leadership for reliability, scalability, and operational excellence across its advertising infrastructure and revenue-critical systems.

Apache Spark ClickHouse GCP Go Kafka Kubernetes
1 day ago

Sr Lead Network Reliability Engineer

Coupa Software 1K-5K Internet Software & Services

Coupa is hiring a Sr. Lead Network Development Engineer to scale and operate its global SaaS platform’s cloud networking infrastructure through automation, reliability engineering, and technical leadership.

Ansible AWS Azure Chef DNS Fortinet Go Java Kubernetes Linux Python Ruby TCP/IP Terraform TLS
2 days ago

Senior Monitoring/Observability Architect

Makpar 51-250 Internet Software & Services

Makpar is seeking a Senior Monitoring/Observability Architect to lead enterprise monitoring strategy, architecture, and implementation guidance for a large federal government program.

Datadog Splunk
2 days ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers