PlayON! Sports Network

PlayON! Sports Network

PlayON! Sports Network provides a comprehensive platform for high school sports programs, offering digital ticketing, live streaming, statistics, coaching tools, and social content to enhance community engagement and support student athletes.

Media
51-250
Founded 2006
$10M raised

Description

  • Assess and improve system visibility by reviewing dashboards, metrics, and logs and closing observability gaps.
  • Tighten monitoring and alerting for critical services to detect issues earlier and improve response times.
  • Build observability into build and deploy workflows by adding instrumentation and telemetry to release processes.
  • Help define SLIs and SLOs for core user flows and align the team on reliability expectations.
  • Improve incident response by partnering with the Event Commander/on-call rotation and strengthening communication, coordination, and follow-up.
  • Automate routine checks and monitoring tasks to reduce manual effort and free up engineering time.
  • Develop automation, tooling, and monitoring solutions that support high service availability.
  • Partner with application and quality engineering teams on reliability practices, release automation, and testing.
  • Drive operational excellence through incident prevention, blameless postmortems, and capacity planning.
  • Participate in on-call rotations to support critical services and respond quickly to incidents.

Requirements

  • Solid experience in Python for automation, tooling, and data-driven operational work.
  • Proficiency in at least one of Java, C++, or Go.
  • Strong understanding of Linux systems, cloud infrastructure, and modern deployment practices.
  • Experience with AWS, GCP, or Azure.
  • Experience with Docker, Kubernetes, and Terraform.
  • Experience with CI/CD pipelines, version control, and automated testing frameworks.
  • Experience with observability tools such as Prometheus, Grafana, ELK, or Datadog.
  • Experience analyzing logs and metrics to diagnose issues.
  • Proven experience facilitating and documenting Critical User Journeys and translating them into actionable SLAs/SLOs for automation.
  • Strong collaboration and communication skills in cross-functional, high-impact situations.
  • Familiarity with AI-augmented development tools such as Claude and Codex.
  • Nice to have: experience writing or maintaining end-to-end or integration tests for distributed systems.
  • Nice to have: background in performance testing, capacity planning, or chaos engineering.
  • Nice to have: contributions to internal developer tooling or reliability-focused frameworks.
  • Nice to have: exposure to security, compliance, or change management processes in production environments.
  • Nice to have: relevant certifications.

Benefits

  • Multiple medical insurance plans to choose from.
  • Dental, vision, life, and disability insurance.
  • Employee Emergency Fund.
  • Company equity in the form of stock options.
  • Open PTO policy.
  • 401(k) plan with company match.
  • Hybrid/flexible work environment.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer (SRE/DevOps)

qode Internet Software & Services

Senior DevOps Engineer at a company building secure, scalable cloud and AI platforms, focused on taking AI systems into production and improving reliability, observability, and operational excellence.

Argo CD AWS Azure CI/CD CloudFormation Flux GCP GitOps NestJS Node.js Pulumi Python Secrets Management Terraform
3 hours, 16 minutes ago

AI infrastructure Engineer (SRE) Bangalore

Together 1-10 IT Services

Together AI is hiring an AI Infrastructure Engineer (SRE) to keep its user-facing services and production systems reliable, scalable, and available as the company builds next-generation AI infrastructure.

Ansible Kubernetes Machine Learning PagerDuty Terraform
2 days, 2 hours ago

Senior Site Reliability Engineer

Omilia 251-1K IT Services

Omilia is hiring a Senior Site Reliability Engineer to operate and improve cloud-based production platforms, observability, and reliability practices across development and engineering teams.

Agile Ansible AWS Bash CentOS Go Grafana Kubernetes MySQL PostgreSQL Prometheus Python Redis TCP/IP Terraform
2 days, 3 hours ago

Site Reliability Engineer II

MRSOOL 1K-5K Air Freight & Logistics

Mrsool is hiring an experienced Site Reliability Engineer to help ensure the stability and reliability of its delivery platform while supporting feature delivery and infrastructure growth.

Ansible AWS Azure Chef Docker GCP Go Grafana Java Kubernetes Nagios Prometheus Puppet Python Ruby Terraform
2 days, 3 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers