Alpaca

Alpaca

Alpaca is a developer-first API for stock and crypto trading, offering easy-to-use APIs for building apps and trading algorithms.

Capital Markets
51-250
Founded 2015
$87M raised

Description

  • Design and evolve cloud architecture on Google Cloud Platform, including networking, interconnects, IAM, and high-availability topology, expressed as code with Terraform.
  • Build and own CI/CD pipelines for infrastructure-as-code changes, including planning, review, testing, policy guardrails, drift detection, and progressive rollout.
  • Develop self-serve platform capabilities and paved paths that let engineers provision infrastructure through a golden-path model.
  • Strengthen observability across metrics, logs, traces, and alerting using tools such as Prometheus, Thanos, Grafana, Loki, Tempo, and Alertmanager.
  • Operate GKE clusters and the workloads running on them, including Helm-packaged services, message brokers, and data stores.
  • Participate in a follow-the-sun on-call rotation to monitor alerts, triage issues, join and declare incidents, and drive debugging and escalation.
  • Lead structured incident response, blameless post-mortems, and follow-up actions that close the loop on operational issues.
  • Embed SRE practices such as SLIs/SLOs, error budgets, and capacity planning into infrastructure operations.
  • Partner closely with SRE and database specialists on the operation of core infrastructure and data services.

Requirements

  • 5+ years of experience in a DevOps, Platform/Infrastructure, or SRE role operating large-scale, high-availability production systems.
  • Deep hands-on experience with Google Cloud Platform as the primary cloud, including landing zones, networking, IAM, and high-availability design.
  • Strong Infrastructure-as-Code experience with Terraform across large multi-environment codebases, using GitOps and least-privilege principles.
  • Proven experience building CI/CD pipelines for infrastructure-as-code, including automated plan/apply, code review, policy-as-code, drift detection, and safe rollout.
  • Significant production experience with Kubernetes, ideally GKE, and deploying workloads with Helm.
  • Strong cloud networking fundamentals across VPCs, routing, load balancing, DNS, TLS, and interconnects.
  • Hands-on experience with observability tooling such as Prometheus, Thanos, Grafana, Loki, Tempo, and Alertmanager.
  • Operator-level familiarity with PostgreSQL and message brokers such as RabbitMQ or RedPanda in production.
  • Understanding of SRE practices, including SLOs, error budgets, capacity planning, and a platform-as-a-product mindset.
  • Strong incident management experience, including debugging under pressure, escalation, documentation, and post-mortems.
  • Willingness to participate in a follow-the-sun on-call rotation from APAC hours and work effectively in an async-first distributed team.
  • Bonus points for experience with OPA/Conftest, Checkov, tflint, Atlantis, Backstage, Tilt, Alloy, Rootly, Go, Linux, Docker/containerd, SOC 2, secrets management, audit logging, trading/brokerage, or low-latency systems.

Benefits

  • Competitive salary with stock options.
  • Health benefits.
  • One-time USD $500 new hire home-office setup stipend.
  • USD $150 monthly stipend via a Brex Card.
  • Opportunity to work on a globally distributed team in a flexible remote environment.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior DevOps Engineer - PA166

ZoomInfo 1K-5K Professional Services

ZoomInfo is hiring a DevOps/Site Reliability Engineer to support its cloud-native data platform by improving infrastructure reliability, scalability, automation, security, and cost efficiency across AWS and GCP environments.

Agile Apache Airflow Apache Spark Argo CD AWS CI/CD Datadog GCP GitOps Grafana Jenkins Kafka Kubernetes Prometheus Python RabbitMQ Snowflake Terraform
17 hours, 21 minutes ago

Pillar Lead: Automation Engineer (R-00207)

True Zero Technologies 11-50 Internet Software & Services

True Zero Technologies is seeking an Automation Engineer to support Zero Trust initiatives by building reusable workflows, integrations, scripts, and automated evidence and reporting mechanisms for Federal cybersecurity programs.

AWS Azure CI/CD Cybersecurity DevSecOps PowerShell Python
18 hours, 6 minutes ago

Sr Data Ops Engineer

Coderio 51-250 Internet Software & Services

Coderio is hiring a DataOps Engineer and Technical Referent to work with international customers on designing, implementing, and operating scalable cloud data infrastructure and automation solutions.

AWS Bash dbt Docker GitHub Python SQL Terraform
18 hours, 6 minutes ago

DevOps Engineer (Cloud) - Freelance

Lingaro 5K-10K IT Services

Build and evolve an Azure-based monitoring and observability platform that improves system reliability, data quality, operational insight, automation, and cloud cost management across supply chain systems.

Ansible Azure Bash CI/CD Databricks Docker GitHub Actions Grafana Kafka Kubernetes Linux Power BI PowerShell Prometheus Python SonarQube Terraform Windows Server
1 day, 17 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers