Alpaca

Alpaca

Alpaca is a developer-first API for stock and crypto trading, offering easy-to-use APIs for building apps and trading algorithms.

Capital Markets
51-250
Founded 2015
$87M raised

Description

  • Design and evolve cloud architecture on Google Cloud Platform, including networking, interconnects, IAM, and high-availability topology, expressed as code with Terraform.
  • Build and own CI/CD pipelines for infrastructure-as-code changes, including planning, review, testing, policy guardrails, drift detection, and progressive rollout.
  • Develop self-serve platform capabilities and paved paths that let engineers provision infrastructure through a golden-path model.
  • Strengthen observability across metrics, logs, traces, and alerting using tools such as Prometheus, Thanos, Grafana, Loki, Tempo, and Alertmanager.
  • Operate GKE clusters and the workloads running on them, including Helm-packaged services, message brokers, and data stores.
  • Participate in a follow-the-sun on-call rotation to monitor alerts, triage issues, join and declare incidents, and drive debugging and escalation.
  • Lead structured incident response, blameless post-mortems, and follow-up actions that close the loop on operational issues.
  • Embed SRE practices such as SLIs/SLOs, error budgets, and capacity planning into infrastructure operations.
  • Partner closely with SRE and database specialists on the operation of core infrastructure and data services.

Requirements

  • 5+ years of experience in a DevOps, Platform/Infrastructure, or SRE role operating large-scale, high-availability production systems.
  • Deep hands-on experience with Google Cloud Platform as the primary cloud, including landing zones, networking, IAM, and high-availability design.
  • Strong Infrastructure-as-Code experience with Terraform across large multi-environment codebases, using GitOps and least-privilege principles.
  • Proven experience building CI/CD pipelines for infrastructure-as-code, including automated plan/apply, code review, policy-as-code, drift detection, and safe rollout.
  • Significant production experience with Kubernetes, ideally GKE, and deploying workloads with Helm.
  • Strong cloud networking fundamentals across VPCs, routing, load balancing, DNS, TLS, and interconnects.
  • Hands-on experience with observability tooling such as Prometheus, Thanos, Grafana, Loki, Tempo, and Alertmanager.
  • Operator-level familiarity with PostgreSQL and message brokers such as RabbitMQ or RedPanda in production.
  • Understanding of SRE practices, including SLOs, error budgets, capacity planning, and a platform-as-a-product mindset.
  • Strong incident management experience, including debugging under pressure, escalation, documentation, and post-mortems.
  • Willingness to participate in a follow-the-sun on-call rotation from APAC hours and work effectively in an async-first distributed team.
  • Bonus points for experience with OPA/Conftest, Checkov, tflint, Atlantis, Backstage, Tilt, Alloy, Rootly, Go, Linux, Docker/containerd, SOC 2, secrets management, audit logging, trading/brokerage, or low-latency systems.

Benefits

  • Competitive salary with stock options.
  • Health benefits.
  • One-time USD $500 new hire home-office setup stipend.
  • USD $150 monthly stipend via a Brex Card.
  • Opportunity to work on a globally distributed team in a flexible remote environment.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Elastic Engineer

Jolera 251-1K Internet Software & Services

Jolera is seeking an Elastic Engineer to design and operate scalable Elasticsearch-based environments for cybersecurity analytics, threat hunting, and detection workflows.

Cybersecurity Elasticsearch Encryption Java Kibana Linux Logstash Machine Learning Python Ruby
15 hours, 31 minutes ago

Senior Fullstack SDK Engineer

Customer.io 11-50 Professional Services

Customer.io is hiring a Fullstack SDK Engineer to build and improve the developer tools and SDKs that power its communication platform across multiple mobile and backend ecosystems.

Android Ember Go iOS Kotlin LLM Microservices Node.js React Swift
15 hours, 31 minutes ago

DevOps Engineer (worldwide remote, work anywhere)

CloudLinux 51-250 IT Services

CloudLinux is hiring a DevOps Engineer for the Patchman team to stabilize, modernize, and integrate infrastructure while improving reliability, performance, and maintainability of the product platform.

Ansible CI/CD ClickHouse Django Docker Git Kafka Kubernetes Linux PostgreSQL Puppet RabbitMQ Redis SaltStack
15 hours, 31 minutes ago

Senior DevOps Automation Engineer

Excella 251-1K Internet Software & Services

Excella is hiring a Senior DevOps Automation Engineer to design and maintain continuous delivery solutions for software delivery teams in a distributed, client-facing environment.

Agile Ansible AWS Azure Bamboo Bash C# Chef CI/CD CircleCI Docker Git GitLab Go Jenkins Kubernetes Node.js Packer PowerShell Puppet Python Ruby Travis CI
16 hours, 1 minute ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers