Pinterest

Pinterest

Pinterest is the world's first visual discovery engine, offering a vast dataset of ideas with over 200 billion recipes, home hacks, and style inspiration. With a mission to inspire everyone to create a life they love, Pinterest empowers its employees t...

Internet Software & Services
5K-10K
Founded 2010

Description

  • Design and build AI agents that support production reliability work, including service health analysis, recommendations, migration playbooks, and risk identification.
  • Lead large-scale infrastructure modernization efforts, including Kubernetes adoption and platform transitions.
  • Transform consulting engagements and reliability patterns into reusable platforms, tools, automation, and self-service documentation.
  • Build the knowledge infrastructure for operational agents, including runbooks, incident patterns, migration playbooks, and best practices.
  • Develop software solutions that improve the reliability and operability of large-scale distributed systems.
  • Create tools, frameworks, and automation that reduce operational toil and overhead.
  • Develop meaningful SLIs that provide actionable signals of system health.
  • Automate critical engineering processes to improve deployment safety and speed at scale.
  • Partner with teams to plan and optimize capacity across public and private cloud environments.

Requirements

  • 5+ years of industry experience building and operating large-scale, high-performance distributed systems.
  • Bachelor's degree in Computer Science or related field, or equivalent experience.
  • Strong programming skills in Python or Go.
  • Deep knowledge of Linux/Unix internals.
  • Experience with open source infrastructure such as MySQL, Kafka, Envoy, or Hadoop.
  • Infrastructure as Code experience with tools such as Terraform, Puppet, Chef, Ansible, Docker, or Kubernetes.
  • Experience deploying web applications to cloud infrastructure such as AWS, GCP, or Azure.
  • Experience working with distributed, service-oriented architecture.
  • Preferred: experience developing AI agents for infrastructure automation, operational decision-making, or reliability workflows.
  • Preferred: AI/ML infrastructure experience, including LLM-based systems, model serving, or agentic workflows.
  • Preferred: technical consulting or embedded SRE experience with cross-functional engineering teams.

Benefits

  • Base salary range of $139,764 to $287,749 USD for US-based applicants.
  • Eligible for equity.
  • Remote-friendly working model with in-office collaboration required only 1-2 times every 6 months.
  • No relocation assistance is provided for this role.
  • Access to Pinterest culture and benefits information via the company benefits page.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Site Reliability Engineer 2 (Azure)

PhonePe 5K-10K Capital Markets

PhonePe Limited is hiring a Site Reliability Engineer to manage and scale core cloud infrastructure for a high-volume digital payments environment in India.

Ansible Azure Bash DNS Docker Go Grafana HAProxy InfluxDB Java Linux MySQL Nginx Prometheus Python RabbitMQ SaltStack Terraform Ubuntu
17 hours, 47 minutes ago

Senior AI Enablement Engineer

Fundraise Up 51-250 Capital Markets

Fundraise Up is hiring an AI Enablement Engineer in a remote role based in Turkey to drive adoption of AI-assisted development and build the internal integrations that improve engineering productivity across a global product team.

Bull CI/CD ClickHouse Elasticsearch Grafana Kafka Koa LLM MongoDB NestJS Node.js OpenTelemetry Prometheus React Redis Serverless TypeScript Vue.js
17 hours, 47 minutes ago

Technical Director, AI Enterprise Architect

Locus Robotics 251-1K Automotive

Locus Robotics is hiring a Technical Director, AI Enterprise Architect to lead an enterprise-wide AI transformation by defining strategy, architecting scalable AI solutions, and driving execution across the business.

AWS Azure CRM Databricks ERP GCP LLM Machine Learning Python
18 hours, 2 minutes ago

Sr. Control System Engineer/Site Reliability Engineer (SRE)

QuEra Computing 11-50 Internet Software & Services

QuEra is seeking a Sr. Control System Engineer/Site Reliability Engineer to integrate and maintain the hardware and software systems that support its quantum control stack and keep development and production environments reliable.

Ansible Bash CI/CD Debian DHCP DNS Docker ELK Stack Embedded Systems Git GitLab CI Go Grafana Jenkins Kubernetes Linux Prometheus Python TCP/IP Terraform Ubuntu
18 hours, 17 minutes ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers