Pinterest

Pinterest

Pinterest is the world's first visual discovery engine, offering a vast dataset of ideas with over 200 billion recipes, home hacks, and style inspiration. With a mission to inspire everyone to create a life they love, Pinterest empowers its employees t...

Internet Software & Services
5K-10K
Founded 2010

Description

  • Ensure the reliability, availability, and performance of production infrastructure and platform services.
  • Operate and scale Kubernetes platforms, including governance and support for multi-tenant workloads.
  • Manage GitOps-based deployment workflows using ArgoCD and Helm.
  • Drive infrastructure provisioning and change management through Terraform and Terragrunt.
  • Build and support CI/CD automation and deployment workflows using GitHub Actions.
  • Lead incident response, root cause analysis, and post-incident improvement efforts.
  • Reduce operational toil through scripting, tooling, and process automation.
  • Advance observability across logs, metrics, traces, dashboards, and alerting.
  • Support secure secrets integration, IAM-aware operations, and platform guardrails.
  • Partner with application, security, and platform teams to improve reliability and delivery outcomes.

Requirements

  • 4+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Cloud Infrastructure.
  • Strong hands-on experience operating AWS in production environments.
  • Deep expertise in Kubernetes, including cluster operations, troubleshooting, workload reliability, and platform administration.
  • Experience with Kubernetes multi-tenancy, including namespaces, RBAC, quotas, policies, and tenant isolation patterns.
  • Experience implementing and operating ArgoCD within a GitOps delivery model.
  • Strong hands-on experience with Helm.
  • Strong experience with Terraform and Terragrunt for infrastructure provisioning and environment management.
  • Solid scripting and automation skills using Bash and/or Python.
  • Experience building, maintaining, or supporting CI/CD pipelines, ideally using GitHub Actions.
  • Strong troubleshooting skills across Linux, containers, IAM, networking, and distributed systems.
  • Experience with monitoring, alerting, and observability in production environments.
  • Demonstrated ownership mindset with experience handling incidents and production issues.
  • Strong collaboration and communication skills across engineering, security, and platform teams.
  • Bachelor’s degree in computer science, engineering, a related field, or equivalent experience.
  • Ability to use AI to improve speed and quality in day-to-day workflow.
  • Ability to critically evaluate and verify AI-assisted work through testing, source-checking, data validation, or peer review.
  • High integrity and ownership, including protecting sensitive data and remaining accountable for final decisions.
  • Relocation assistance is not available for this role.

Benefits

  • Base salary range of $139,764 to $287,749 USD for US-based applicants.
  • Eligible for equity.
  • Flexible PinFlex working model.
  • Information about Pinterest culture and benefits is available to candidates.
  • Remote work designation noted with #LI-REMOTE.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Site Reliability Engineering (SRE) Leader

PatSnap 251-1K Internet Software & Services

PatSnap is hiring a Site Reliability Engineering (SRE) Leader to lead its UK SRE team and drive the reliability, scalability, security, and performance of its global SaaS platform.

AWS Docker Kubernetes
23 hours, 26 minutes ago

SITE RELIABILITY ENGINEER III

Harford County Public Library 51-250 Diversified Consumer Services

Site Reliability Engineer na Stone, atuando no time de Foundation Platform para fortalecer a plataforma interna de tecnologia com foco em observabilidade, automação e estabilidade dos sistemas.

Ansible Argo CD AWS Azure Datadog Docker GCP GitHub Actions Go Grafana Kubernetes Linux Node.js OpenTelemetry Prometheus Python Splunk Terraform
1 day, 22 hours ago

Sr. Site Reliability Engineer (Starlink)

SpaceX 10K-50K Aerospace & Defense

SpaceX is hiring a Sr. Site Reliability Engineer for Starlink to improve the reliability, scalability, and performance of the systems supporting its satellite internet service.

Apache Spark C# CI/CD Flink Git Go HDFS Java Kafka Kubernetes Linux Python Scala
1 day, 22 hours ago

Head of Platform Engineering

dLocal 251-1K Diversified Financial Services

dLocal is seeking a senior leader to own its engineering platform, reliability posture, and AI-assisted development transformation across a global payments business serving emerging markets.

CI/CD Microservices
1 day, 22 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers