PENN Entertainment

PENN Entertainment

PENN Entertainment is a leading provider of integrated entertainment, sports content, and casino gaming experiences across North America, offering a diverse range of entertainment destinations and a premier loyalty program that rewards members with cas...

Hotels, Restaurants & Leisure
10K-50K
Founded 1982

Description

  • Lead real-time incident management across multiple incident severity levels, including P1 through P4.
  • Classify, document, investigate, and drive resolution of incidents through diagnosis, recovery, and root cause analysis.
  • Coordinate with Incident Commanders, Customer Support, Application, Engineering, SRE, and Infrastructure teams during incidents.
  • Improve service delivery and release processes based on disruption reports.
  • Develop and maintain practices, frameworks, process flows, templates, and process guides for incident management.
  • Continuously improve internal frameworks, methodologies, processes, and tools.
  • Identify requirements and gaps with SRE and Infrastructure teams that may cause downtime or blind spots.
  • Lead stakeholder communications through email, Slack, and Teams in a timely manner.
  • Promote JIRA release ticket management and alignment with incident communication and SLA requirements.
  • Maintain root cause analysis documentation and support continuous service improvement initiatives.

Requirements

  • Experience in a similar role or incident management role.
  • Experience and understanding of containerization, with Docker and Kubernetes preferred.
  • Experience with configuration management and infrastructure as code tools such as Terraform, Ansible, Helm, or similar.
  • Experience with a programming language.
  • Comfortable working in Linux environments.
  • Experience working across AWS, GCP, and on-premises environments.
  • Ability to work independently, learn quickly, and handle multiple projects simultaneously.
  • Willingness to respond to ad hoc tasks and shift priorities quickly.
  • Strong communication skills and an outgoing, conversational approach to gathering information.
  • A degree in computer science, engineering, or similar experience.
  • Nice to have: Postgres, MySQL, Elasticsearch, Kafka, Redis, Terragrunt, Prometheus, Python, and Talos Linux.

Benefits

  • Competitive compensation package with a $90,000 to $135,000 USD salary range.
  • Fun, relaxed work environment.
  • Education and conference reimbursements.
  • Opportunities for career progression and mentoring others.
  • Bonus eligibility for most non-sales positions.
  • Best-in-class benefits for eligible employees.
  • Remote work designation (#LI-REMOTE).

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

GoDaddy 5001-10000 Technology, Information and Internet

GoDaddy is hiring a Senior Site Reliability Engineer to support the Commerce ecosystem remotely by improving the reliability, scalability, security, and operation of business-critical production platforms.

Ansible AWS AWS CDK CI/CD CloudFormation Go Kubernetes Linux Pulumi Python SaltStack Terraform TypeScript
57 minutes ago

Staff Site Reliability Engineer, Ads

Reddit 1K-5K Internet Software & Services

Reddit is hiring a Staff Site Reliability Engineer to provide technical leadership for reliability, scalability, and operational excellence across its advertising infrastructure and revenue-critical systems.

Apache Spark ClickHouse GCP Go Kafka Kubernetes
1 day ago

Sr Lead Network Reliability Engineer

Coupa Software 1K-5K Internet Software & Services

Coupa is hiring a Sr. Lead Network Development Engineer to scale and operate its global SaaS platform’s cloud networking infrastructure through automation, reliability engineering, and technical leadership.

Ansible AWS Azure Chef DNS Fortinet Go Java Kubernetes Linux Python Ruby TCP/IP Terraform TLS
2 days ago

Senior Monitoring/Observability Architect

Makpar 51-250 Internet Software & Services

Makpar is seeking a Senior Monitoring/Observability Architect to lead enterprise monitoring strategy, architecture, and implementation guidance for a large federal government program.

Datadog Splunk
2 days ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers