Pinterest

Pinterest

Pinterest is the world's first visual discovery engine, offering a vast dataset of ideas with over 200 billion recipes, home hacks, and style inspiration. With a mission to inspire everyone to create a life they love, Pinterest empowers its employees t...

Internet Software & Services
5K-10K
Founded 2010

Description

  • Ensure the reliability, availability, and performance of production infrastructure and platform services.
  • Operate and scale Kubernetes platforms, including governance and support for multi-tenant workloads.
  • Manage GitOps-based deployment workflows using ArgoCD and Helm.
  • Drive infrastructure provisioning and change management through Terraform and Terragrunt.
  • Build and support CI/CD automation and deployment workflows using GitHub Actions.
  • Lead incident response, root cause analysis, and post-incident improvement efforts.
  • Reduce operational toil through scripting, tooling, and process automation.
  • Advance observability across logs, metrics, traces, dashboards, and alerting.
  • Support secure secrets integration, IAM-aware operations, and platform guardrails.
  • Partner with application, security, and platform teams to improve reliability and delivery outcomes.

Requirements

  • 4+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Cloud Infrastructure.
  • Strong hands-on experience operating AWS in production environments.
  • Deep expertise in Kubernetes, including cluster operations, troubleshooting, workload reliability, and platform administration.
  • Experience with Kubernetes multi-tenancy, including namespaces, RBAC, quotas, policies, and tenant isolation patterns.
  • Experience implementing and operating ArgoCD within a GitOps delivery model.
  • Strong hands-on experience with Helm.
  • Strong experience with Terraform and Terragrunt for infrastructure provisioning and environment management.
  • Solid scripting and automation skills using Bash and/or Python.
  • Experience building, maintaining, or supporting CI/CD pipelines, ideally using GitHub Actions.
  • Strong troubleshooting skills across Linux, containers, IAM, networking, and distributed systems.
  • Experience with monitoring, alerting, and observability in production environments.
  • Demonstrated ownership mindset with experience handling incidents and production issues.
  • Strong collaboration and communication skills across engineering, security, and platform teams.
  • Bachelor’s degree in computer science, engineering, a related field, or equivalent experience.
  • Ability to use AI to improve speed and quality in day-to-day workflow.
  • Ability to critically evaluate and verify AI-assisted work through testing, source-checking, data validation, or peer review.
  • High integrity and ownership, including protecting sensitive data and remaining accountable for final decisions.
  • Relocation assistance is not available for this role.

Benefits

  • Base salary range of $139,764 to $287,749 USD for US-based applicants.
  • Eligible for equity.
  • Flexible PinFlex working model.
  • Information about Pinterest culture and benefits is available to candidates.
  • Remote work designation noted with #LI-REMOTE.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Technical Support Engineer (GPU Clusters) - US Weekends

Together 1-10 IT Services

Together AI is hiring a Technical Support Engineer to support customers building and operating AI training, fine-tuning, and inference systems on Kubernetes GPU infrastructure.

Ansible Kubernetes Machine Learning
9 hours, 48 minutes ago

Staff Site Reliability Engineer

BeyondTrust 1K-5K Professional Services

BeyondTrust is hiring a Staff Site Reliability Engineer to lead the evolution of its Password Safe platform, infrastructure, and deployment ecosystem across cloud and on-premises environments.

Ansible AWS Azure C# CI/CD Datadog DevSecOps Docker GitOps Go Java Kubernetes Linux Microservices OpenTelemetry Secrets Management Terraform
9 hours, 48 minutes ago

Senior Site Reliability Engineer

Alpaca 51-250 Capital Markets

Alpaca is hiring a Site Reliability Engineer to keep its brokerage platform reliable, observable, and operable across cloud infrastructure, Kubernetes, and PostgreSQL on the trading-critical path.

DNS GitOps Go Kafka Kubernetes Linux Load Balancing PostgreSQL Python RabbitMQ Secrets Management TLS
1 day, 9 hours ago

Software/Site Reliability Engineer - FedRAMP

Tenable 1K-5K Internet Software & Services

Tenable is hiring a Site Reliability Engineer to help scale and operate its cloud-based vulnerability management platform for private and U.S. Government cloud customers.

Agile AWS Azure Bash CI/CD Datadog Docker DynamoDB Elasticsearch GCP Go Gradle Groovy Helm Java Kafka Kotlin Kubernetes Microservices Node.js OpenSearch OpenTelemetry Python Splunk Terraform
3 days, 9 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers