Alpaca

Alpaca

Alpaca is a developer-first API for stock and crypto trading, offering easy-to-use APIs for building apps and trading algorithms.

Capital Markets
51-250
Founded 2015
$87M raised

Description

  • Design and evolve cloud architecture on Google Cloud Platform, including networking, interconnects, IAM, and high-availability topology, expressed as code with Terraform.
  • Build and own CI/CD pipelines for infrastructure-as-code changes, including planning, review, testing, policy guardrails, drift detection, and progressive rollout.
  • Develop self-serve platform capabilities and paved paths that let engineers provision infrastructure through a golden-path model.
  • Strengthen observability across metrics, logs, traces, and alerting using tools such as Prometheus, Thanos, Grafana, Loki, Tempo, and Alertmanager.
  • Operate GKE clusters and the workloads running on them, including Helm-packaged services, message brokers, and data stores.
  • Participate in a follow-the-sun on-call rotation to monitor alerts, triage issues, join and declare incidents, and drive debugging and escalation.
  • Lead structured incident response, blameless post-mortems, and follow-up actions that close the loop on operational issues.
  • Embed SRE practices such as SLIs/SLOs, error budgets, and capacity planning into infrastructure operations.
  • Partner closely with SRE and database specialists on the operation of core infrastructure and data services.

Requirements

  • 5+ years of experience in a DevOps, Platform/Infrastructure, or SRE role operating large-scale, high-availability production systems.
  • Deep hands-on experience with Google Cloud Platform as the primary cloud, including landing zones, networking, IAM, and high-availability design.
  • Strong Infrastructure-as-Code experience with Terraform across large multi-environment codebases, using GitOps and least-privilege principles.
  • Proven experience building CI/CD pipelines for infrastructure-as-code, including automated plan/apply, code review, policy-as-code, drift detection, and safe rollout.
  • Significant production experience with Kubernetes, ideally GKE, and deploying workloads with Helm.
  • Strong cloud networking fundamentals across VPCs, routing, load balancing, DNS, TLS, and interconnects.
  • Hands-on experience with observability tooling such as Prometheus, Thanos, Grafana, Loki, Tempo, and Alertmanager.
  • Operator-level familiarity with PostgreSQL and message brokers such as RabbitMQ or RedPanda in production.
  • Understanding of SRE practices, including SLOs, error budgets, capacity planning, and a platform-as-a-product mindset.
  • Strong incident management experience, including debugging under pressure, escalation, documentation, and post-mortems.
  • Willingness to participate in a follow-the-sun on-call rotation from APAC hours and work effectively in an async-first distributed team.
  • Bonus points for experience with OPA/Conftest, Checkov, tflint, Atlantis, Backstage, Tilt, Alloy, Rootly, Go, Linux, Docker/containerd, SOC 2, secrets management, audit logging, trading/brokerage, or low-latency systems.

Benefits

  • Competitive salary with stock options.
  • Health benefits.
  • One-time USD $500 new hire home-office setup stipend.
  • USD $150 monthly stipend via a Brex Card.
  • Opportunity to work on a globally distributed team in a flexible remote environment.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

PnP Pipeline On-call Support Engineer

PHIZENIX 11-50 information technology & services

This role at an internal engineering organization focuses on monitoring and triaging PnP (Power and Performance) pipelines to keep post-silicon work stable and moving efficiently.

CI/CD Jest
16 hours, 7 minutes ago

Senior Full Stack Engineer

Virtru 51-250 IT Services

Virtru is hiring a Senior Full Stack Engineer to build and operate its digital privacy products and core data protection platform for organizations that manage sensitive data.

AWS Docker GCP Go JavaScript Node.js PagerDuty React Selenium
2 days, 15 hours ago

Staff Forward Deployed Engineer

Tenstorrent 251-1K Internet Software & Services

Tenstorrent is hiring a remote Forward Deployed Engineer in North America to work directly with customers and internal teams on production AI inference deployments for its AI computers.

Grafana Helm Kubernetes OpenTelemetry Prometheus
2 days, 15 hours ago

Senior DevOps / MLOps Engineer (AI Agents, Claude)

Xebia 1K-5K Internet Software & Services

Xebia is hiring DevOps/MLOps Engineers for an international AI platform project focused on building the operational foundations for autonomous AI agent applications in the .NET ecosystem.

Azure CI/CD Docker Generative AI Linux LLM MLOps .NET Secrets Management
2 days, 16 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers