Alpaca

Alpaca

Alpaca is a developer-first API for stock and crypto trading, offering easy-to-use APIs for building apps and trading algorithms.

Capital Markets
51-250
Founded 2015
$87M raised

Description

  • Operate production systems day to day, including on-call support, incident response, and postmortems.
  • Define and refine SLIs, SLOs, and error budgets to improve reliability.
  • Improve observability across metrics, logs, traces, and alerting.
  • Ship infrastructure as code through a GitOps workflow for cloud resources and Kubernetes workloads.
  • Support PostgreSQL reliability through performance tuning, schema and migration review, online migrations, high availability/disaster recovery, and CDC pipelines.
  • Mentor engineers on reliability and database fundamentals through code review, design review, and pairing.

Requirements

  • 4+ years of experience in SRE, DevOps, Platform/Infrastructure, or backend engineering with significant production operations ownership.
  • Hands-on experience operating production services on Kubernetes and shipping infrastructure as code in a GitOps workflow.
  • Solid working knowledge of PostgreSQL in production, including query plans, pg_stat_*, indexing, schema trade-offs, and safe online migrations on non-trivial tables.
  • Cloud networking fundamentals, including VPCs, routing, L4/L7 load balancing, DNS, and TLS, with comfort debugging cross-service connectivity.
  • Comfort with a modern observability stack and proficiency with Linux at the operator level.
  • Experience with incident response, including structured debugging and postmortems that drive change.
  • Working proficiency in Go or Python, plus strong written and verbal communication skills.
  • Genuine interest in databases and in growing PostgreSQL/DBA expertise.
  • Deeper PostgreSQL experience with large OLTP clusters, online migrations on big tables, HA/DR ownership, connection pooling at scale, or change-data-capture pipelines (preferred).
  • Experience with typed SQL access layers in Go such as pgx, gorm, or sqlc (preferred).
  • Production experience with messaging systems at scale such as RabbitMQ, Kafka, or Redpanda (preferred).
  • Security and compliance experience in a regulated environment, including SOC 2, secrets management, or audit logging (preferred).
  • Familiarity with trading, brokerage, or other regulated fintech domains (preferred).

Benefits

  • Competitive salary with stock options.
  • Health benefits.
  • One-time USD $500 new hire home-office setup stipend.
  • USD $150 monthly stipend via Brex card.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Software/Site Reliability Engineer - FedRAMP

Tenable 1K-5K Internet Software & Services

Tenable is hiring a Site Reliability Engineer to help scale and operate its cloud-based vulnerability management platform for private and U.S. Government cloud customers.

Agile AWS Azure Bash CI/CD Datadog Docker DynamoDB Elasticsearch GCP Go Gradle Groovy Helm Java Kafka Kotlin Kubernetes Microservices Node.js OpenSearch OpenTelemetry Python Splunk Terraform
2 days, 5 hours ago

Senior Site Reliability Engineer (SRE)

Branch 51-250 Professional Services

Branch is hiring a Senior Site Reliability Engineer to improve the reliability, scalability, performance, and observability of its fintech platform through automation and operational best practices.

Bash Docker GCP Go Gradle Grafana Java Kubernetes MySQL OpenTelemetry Prometheus Python Redis Spring Boot Terraform
2 days, 6 hours ago

Staff Site Reliability Engineer

Filevine 251-1K Specialized Consumer Services

Filevine is seeking a Staff Site Reliability Engineer to lead reliability strategy and production excellence for its cloud platform and distributed systems supporting modern legal operations.

Bash Datadog Go HIPAA Kubernetes Machine Learning New Relic Python
3 days, 7 hours ago

Senior Site Reliability Engineer – Telephony & Communications Platform (AWS)

Filevine 251-1K Specialized Consumer Services

Filevine is hiring a Site Reliability Engineer to strengthen the reliability, scalability, and recoverability of its legal AI platform and the systems that support it.

AWS Azure Bash CI/CD EC2 GCP HIPAA PowerShell Python Terraform Twilio
3 days, 7 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers