PENN Entertainment

PENN Entertainment

PENN Entertainment is a leading provider of integrated entertainment, sports content, and casino gaming experiences across North America, offering a diverse range of entertainment destinations and a premier loyalty program that rewards members with cas...

Hotels, Restaurants & Leisure
10K-50K
Founded 1982

Description

  • Lead real-time incident management across multiple incident severity levels, including P1 through P4.
  • Classify, document, investigate, and drive resolution of incidents through diagnosis, recovery, and root cause analysis.
  • Coordinate with Incident Commanders, Customer Support, Application, Engineering, SRE, and Infrastructure teams during incidents.
  • Improve service delivery and release processes based on disruption reports.
  • Develop and maintain practices, frameworks, process flows, templates, and process guides for incident management.
  • Continuously improve internal frameworks, methodologies, processes, and tools.
  • Identify requirements and gaps with SRE and Infrastructure teams that may cause downtime or blind spots.
  • Lead stakeholder communications through email, Slack, and Teams in a timely manner.
  • Promote JIRA release ticket management and alignment with incident communication and SLA requirements.
  • Maintain root cause analysis documentation and support continuous service improvement initiatives.

Requirements

  • Experience in a similar role or incident management role.
  • Experience and understanding of containerization, with Docker and Kubernetes preferred.
  • Experience with configuration management and infrastructure as code tools such as Terraform, Ansible, Helm, or similar.
  • Experience with a programming language.
  • Comfortable working in Linux environments.
  • Experience working across AWS, GCP, and on-premises environments.
  • Ability to work independently, learn quickly, and handle multiple projects simultaneously.
  • Willingness to respond to ad hoc tasks and shift priorities quickly.
  • Strong communication skills and an outgoing, conversational approach to gathering information.
  • A degree in computer science, engineering, or similar experience.
  • Nice to have: Postgres, MySQL, Elasticsearch, Kafka, Redis, Terragrunt, Prometheus, Python, and Talos Linux.

Benefits

  • Competitive compensation package with a $90,000 to $135,000 USD salary range.
  • Fun, relaxed work environment.
  • Education and conference reimbursements.
  • Opportunities for career progression and mentoring others.
  • Bonus eligibility for most non-sales positions.
  • Best-in-class benefits for eligible employees.
  • Remote work designation (#LI-REMOTE).

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Site Reliability Engineer

MyFitnessPal 10K-50K Health Care Providers & Services

MyFitnessPal is hiring a Software Engineer III, Site Reliability to own production reliability and delivery pipeline security for the PEAS team supporting automation, CI/CD, and self-service platforms.

AWS CI/CD Datadog GitHub Actions Go HIPAA Kubernetes Python Secrets Management Terraform TypeScript
1 hour, 53 minutes ago

Site Reliability Engineer 2 (Azure)

PhonePe 5K-10K Capital Markets

PhonePe Limited is hiring a Site Reliability Engineer to manage and scale core cloud infrastructure for a high-volume digital payments environment in India.

Ansible Azure Bash DNS Docker Go Grafana HAProxy InfluxDB Java Linux MySQL Nginx Prometheus Python RabbitMQ SaltStack Terraform Ubuntu
1 day ago

Sr. Control System Engineer/Site Reliability Engineer (SRE)

QuEra Computing 11-50 Internet Software & Services

QuEra is seeking a Sr. Control System Engineer/Site Reliability Engineer to integrate and maintain the hardware and software systems that support its quantum control stack and keep development and production environments reliable.

Ansible Bash CI/CD Debian DHCP DNS Docker ELK Stack Embedded Systems Git GitLab CI Go Grafana Jenkins Kubernetes Linux Prometheus Python TCP/IP Terraform Ubuntu
1 day, 1 hour ago

Site Reliability Engineer

VantageScore 11-50 Banks

Site Reliability Engineer at a growing engineering team, focused on DevSecOps for maintaining the reliability, security, and compliance of cloud infrastructure, APIs, and software supply chains.

Agile AWS AWS CDK Bash CI/CD CloudFormation CodePipeline Datadog DevSecOps Docker EC2 GitHub Actions Grafana HashiCorp Vault Kong Kubernetes Microservices Python REST API Scrum Terraform
2 days, 1 hour ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers