Pinterest

Pinterest

Pinterest is the world's first visual discovery engine, offering a vast dataset of ideas with over 200 billion recipes, home hacks, and style inspiration. With a mission to inspire everyone to create a life they love, Pinterest empowers its employees t...

Internet Software & Services
5K-10K
Founded 2010

Description

  • Ensure the reliability, availability, and performance of production infrastructure and platform services.
  • Operate and scale Kubernetes platforms, including governance and support for multi-tenant workloads.
  • Manage GitOps-based deployment workflows using ArgoCD and Helm.
  • Drive infrastructure provisioning and change management through Terraform and Terragrunt.
  • Build and support CI/CD automation and deployment workflows using GitHub Actions.
  • Lead incident response, root cause analysis, and post-incident improvement efforts.
  • Reduce operational toil through scripting, tooling, and process automation.
  • Advance observability across logs, metrics, traces, dashboards, and alerting.
  • Support secure secrets integration, IAM-aware operations, and platform guardrails.
  • Partner with application, security, and platform teams to improve reliability and delivery outcomes.

Requirements

  • 4+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Cloud Infrastructure.
  • Strong hands-on experience operating AWS in production environments.
  • Deep expertise in Kubernetes, including cluster operations, troubleshooting, workload reliability, and platform administration.
  • Experience with Kubernetes multi-tenancy, including namespaces, RBAC, quotas, policies, and tenant isolation patterns.
  • Experience implementing and operating ArgoCD within a GitOps delivery model.
  • Strong hands-on experience with Helm.
  • Strong experience with Terraform and Terragrunt for infrastructure provisioning and environment management.
  • Solid scripting and automation skills using Bash and/or Python.
  • Experience building, maintaining, or supporting CI/CD pipelines, ideally using GitHub Actions.
  • Strong troubleshooting skills across Linux, containers, IAM, networking, and distributed systems.
  • Experience with monitoring, alerting, and observability in production environments.
  • Demonstrated ownership mindset with experience handling incidents and production issues.
  • Strong collaboration and communication skills across engineering, security, and platform teams.
  • Bachelor’s degree in computer science, engineering, a related field, or equivalent experience.
  • Ability to use AI to improve speed and quality in day-to-day workflow.
  • Ability to critically evaluate and verify AI-assisted work through testing, source-checking, data validation, or peer review.
  • High integrity and ownership, including protecting sensitive data and remaining accountable for final decisions.
  • Relocation assistance is not available for this role.

Benefits

  • Base salary range of $139,764 to $287,749 USD for US-based applicants.
  • Eligible for equity.
  • Flexible PinFlex working model.
  • Information about Pinterest culture and benefits is available to candidates.
  • Remote work designation noted with #LI-REMOTE.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Site Reliability Engineer

AssureSoft 51-250 Internet Software & Services

AssureSoft is hiring a remote Site Reliability Engineer to support production cloud infrastructure and platform reliability for long-term client projects.

Argo CD AWS Bash DNS Docker GCP GitHub Actions Go Grafana HTTP Kafka Kubernetes Linux Load Balancing Prometheus Python RabbitMQ Snowflake TCP/IP TLS TypeScript Unix
1 day, 9 hours ago

Senior Site Reliability Engineer

Megaport 251-1K Diversified Telecommunication Services

Megaport is hiring a Senior Platform Engineer to support secure, reliable, and maintainable global production systems within its SRE-focused platform team.

AWS Bash Cassandra CI/CD ClickHouse Git GitHub Go Kubernetes Linux PostgreSQL Python Terraform
1 day, 10 hours ago

Site Reliability Engineer - Azure, Observability and Scripting

Jalasoft 1K-5K Internet Software & Services

Jalasoft is hiring a Site Reliability Engineer to support the reliability, scalability, and performance of cloud-native platforms on Microsoft Azure and Kubernetes.

Azure Bash Grafana Kubernetes OpenTelemetry Oracle PowerShell Prometheus Python Terraform
2 days, 10 hours ago

Senior Service Reliability Engineer

Thoughtworks 10K-50K Professional Services

Senior Service Reliability Engineer at Thoughtworks, focused on improving infrastructure reliability, observability, and incident response for customer-facing production systems.

Azure Bash Datadog ELK Stack GitOps Go Grafana Kubernetes Microservices Network Security New Relic Nomad Python REST API Serverless Terraform
2 days, 10 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers