Alpaca

Alpaca

Alpaca is a developer-first API for stock and crypto trading, offering easy-to-use APIs for building apps and trading algorithms.

Capital Markets
51-250
Founded 2015
$87M raised

Description

  • Design and evolve cloud architecture on Google Cloud Platform, including networking, interconnects, IAM, and high-availability topology, expressed as code with Terraform.
  • Build and own CI/CD pipelines for infrastructure-as-code changes, including planning, review, testing, policy guardrails, drift detection, and progressive rollout.
  • Develop self-serve platform capabilities and paved paths that let engineers provision infrastructure through a golden-path model.
  • Strengthen observability across metrics, logs, traces, and alerting using tools such as Prometheus, Thanos, Grafana, Loki, Tempo, and Alertmanager.
  • Operate GKE clusters and the workloads running on them, including Helm-packaged services, message brokers, and data stores.
  • Participate in a follow-the-sun on-call rotation to monitor alerts, triage issues, join and declare incidents, and drive debugging and escalation.
  • Lead structured incident response, blameless post-mortems, and follow-up actions that close the loop on operational issues.
  • Embed SRE practices such as SLIs/SLOs, error budgets, and capacity planning into infrastructure operations.
  • Partner closely with SRE and database specialists on the operation of core infrastructure and data services.

Requirements

  • 5+ years of experience in a DevOps, Platform/Infrastructure, or SRE role operating large-scale, high-availability production systems.
  • Deep hands-on experience with Google Cloud Platform as the primary cloud, including landing zones, networking, IAM, and high-availability design.
  • Strong Infrastructure-as-Code experience with Terraform across large multi-environment codebases, using GitOps and least-privilege principles.
  • Proven experience building CI/CD pipelines for infrastructure-as-code, including automated plan/apply, code review, policy-as-code, drift detection, and safe rollout.
  • Significant production experience with Kubernetes, ideally GKE, and deploying workloads with Helm.
  • Strong cloud networking fundamentals across VPCs, routing, load balancing, DNS, TLS, and interconnects.
  • Hands-on experience with observability tooling such as Prometheus, Thanos, Grafana, Loki, Tempo, and Alertmanager.
  • Operator-level familiarity with PostgreSQL and message brokers such as RabbitMQ or RedPanda in production.
  • Understanding of SRE practices, including SLOs, error budgets, capacity planning, and a platform-as-a-product mindset.
  • Strong incident management experience, including debugging under pressure, escalation, documentation, and post-mortems.
  • Willingness to participate in a follow-the-sun on-call rotation from APAC hours and work effectively in an async-first distributed team.
  • Bonus points for experience with OPA/Conftest, Checkov, tflint, Atlantis, Backstage, Tilt, Alloy, Rootly, Go, Linux, Docker/containerd, SOC 2, secrets management, audit logging, trading/brokerage, or low-latency systems.

Benefits

  • Competitive salary with stock options.
  • Health benefits.
  • One-time USD $500 new hire home-office setup stipend.
  • USD $150 monthly stipend via a Brex Card.
  • Opportunity to work on a globally distributed team in a flexible remote environment.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Developer Advocate - Service Management EMEA

Datadog 5K-10K IT Services

Datadog is seeking a service-management and technical advocacy engineer to help SRE, DevOps, and operations communities improve incident response, observability, and operational automation through engineering and technical storytelling.

Bash Datadog Go Node.js OpsGenie PagerDuty Python
27 minutes ago

[Job-31795] Senior/Specialist DevOps (preferably Azure), Brazil

CI&T 5K-10K Internet Software & Services

CI&T is seeking a Senior/Specialist DevOps Engineer to design, build, and optimize scalable, secure, and reliable platform infrastructure for enterprise technology transformation projects.

AWS Azure Bash CI/CD CircleCI Docker GCP Jenkins Kubernetes Python Terraform
1 hour, 42 minutes ago

Integration & Automation Engineer Sr (Python/TypeScript, n8n)

MUTT DATA 51-250 Internet Software & Services

Muttdata is hiring a Senior Integration & Automation Engineer for its remote-first low-code team, supporting a leading Argentine bank by evaluating, deploying, operating, and extending automation and agent-building platforms.

AWS Azure CI/CD Databricks dbt GCP GitHub Actions Grafana JavaScript OAuth Python REST API Terraform TypeScript
1 hour, 42 minutes ago

DevOps / SRE / DevSecOps Engineer (AWS) - Latin America, Remote

Bluelight Consulting 11-50 Internet Software & Services

Bluelight is hiring a DevOps/SRE/DevSecOps professional to operate secure, reliable AWS-based systems and delivery infrastructure for client software projects in a remote consultancy environment.

AWS CI/CD DevSecOps Docker K6 OpenTelemetry PostgreSQL Secrets Management Terraform
1 day, 1 hour ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers