MyFitnessPal

MyFitnessPal

MyFitnessPal is a top nutrition tracking app with a calorie tracker, BMR calculator, and food tracker. It helps users reach health goals by tracking meals and physical activity, offering nonstop motivation for a healthier life.

Health Care Providers & Services
10K-50K
Founded 2005

Description

  • Own and evolve SLI/SLO and error-budget frameworks to inform prioritization and product decisions.
  • Lead incident response, drive postmortems, and implement systemic fixes.
  • Build and maintain observability across metrics, logs, and traces using Datadog.
  • Design and operate resilient, scalable infrastructure using Terraform.
  • Manage production Kubernetes and container workloads, including capacity planning and cloud-cost optimization.
  • Own CI/CD pipelines and safe deployment strategies such as canary, progressive rollout, and fast rollback.
  • Integrate and tune security scanning in the delivery pipeline, including SAST, DAST, and SCA.
  • Implement and maintain policy-as-code controls to block unsafe infrastructure and Kubernetes changes at admission time.
  • Drive vulnerability triage and remediation SLAs for pipeline and infrastructure findings.
  • Improve on-call operations by building sustainable runbooks and automation, and coach engineers on reliability best practices.

Requirements

  • 5+ years of experience in site reliability, platform, or infrastructure engineering with senior-level ownership of production systems.
  • Strong programming skills for automation and tooling in Go, Python, TypeScript, or similar languages.
  • Deep hands-on experience with a major cloud platform, Kubernetes, and Infrastructure as Code.
  • Experience with AWS is a plus.
  • Experience with Terraform is a plus.
  • Proven track record leading incident response and building SLO-driven reliability practices.
  • Working fluency with observability tooling, with Datadog as a plus.
  • Practical experience integrating security into CI/CD pipelines, including SAST/DAST/SCA, dependency scanning, or policy-as-code.
  • Strong understanding of cloud security fundamentals, including IAM, least privilege, policy guardrails, and secrets management.
  • Experience with Kyverno, OPA/Rego, or Conftest enforced at admission time is a plus.
  • Exposure to regulated or compliance-driven environments such as SOC 2, PCI DSS, or HIPAA is a plus.
  • Chaos engineering or game-day experience is a plus.
  • Experience supporting B2C or mobile backend environments with high traffic and strong reliability needs is a plus.

Benefits

  • Salary range of $120,000-$165,000.
  • Annual performance bonus.
  • Comprehensive healthcare benefits, including medical, dental, and vision.
  • Parental planning support, including paid maternity and paternity leave and fertility assistance.
  • 401(k) retirement plan with employer match.
  • Responsible time off policy.
  • Monthly wellness and technology allowances.
  • Mental health benefits and dedicated mental health days.
  • Access to MyFitnessPal Premium.
  • Learning and development resources and training opportunities.
  • Volunteer days off.
  • Flexible, in-person team connection opportunities and annual company gatherings.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Site Reliability Engineer 2 (Azure)

PhonePe 5K-10K Capital Markets

PhonePe Limited is hiring a Site Reliability Engineer to manage and scale core cloud infrastructure for a high-volume digital payments environment in India.

Ansible Azure Bash DNS Docker Go Grafana HAProxy InfluxDB Java Linux MySQL Nginx Prometheus Python RabbitMQ SaltStack Terraform Ubuntu
23 hours, 44 minutes ago

Sr. Control System Engineer/Site Reliability Engineer (SRE)

QuEra Computing 11-50 Internet Software & Services

QuEra is seeking a Sr. Control System Engineer/Site Reliability Engineer to integrate and maintain the hardware and software systems that support its quantum control stack and keep development and production environments reliable.

Ansible Bash CI/CD Debian DHCP DNS Docker ELK Stack Embedded Systems Git GitLab CI Go Grafana Jenkins Kubernetes Linux Prometheus Python TCP/IP Terraform Ubuntu
1 day ago

Incident Commander

PENN Entertainment 10K-50K Hotels, Restaurants & Leisure

PENN Interactive is hiring an Incident Commander to join its site reliability team and lead cross-functional incident response for its online and physical platforms.

Ansible AWS Docker Elasticsearch GCP Helm JIRA Kafka Kubernetes Linux MySQL PostgreSQL Prometheus Python Redis Terraform
1 day ago

Site Reliability Engineer

VantageScore 11-50 Banks

Site Reliability Engineer at a growing engineering team, focused on DevSecOps for maintaining the reliability, security, and compliance of cloud infrastructure, APIs, and software supply chains.

Agile AWS AWS CDK Bash CI/CD CloudFormation CodePipeline Datadog DevSecOps Docker EC2 GitHub Actions Grafana HashiCorp Vault Kong Kubernetes Microservices Python REST API Scrum Terraform
2 days ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers