Tinybird

Tinybird

Tinybird is a real-time analytics platform that enables data teams and developers to build low-latency APIs in minutes using SQL. It ingests millions of rows per second and serves high-concurrency analytical queries, helping businesses turn data into r...

IT Services
11-50
Founded 2019
$40M raised

Description

  • Build, operate, and continually improve the technical foundations Tinybird relies on.
  • Design and manage distributed cloud architectures and large-scale production systems.
  • Run and optimize Kubernetes infrastructure, including production clusters, autoscaling, and safe deployments.
  • Improve observability through telemetry, dashboards, alerting, and long-term system health visibility.
  • Investigate and resolve production incidents, strengthen disaster recovery, and improve on-call practices.
  • Analyze and improve performance across storage, networking, compute, and ClickHouse.
  • Reduce operational burden by turning manual or fragile processes into repeatable systems.
  • Strengthen CI/CD foundations and support safer, more confident releases.
  • Collaborate with product and backend teams on architecture, resource optimization, and platform evolution.
  • Support customer and internal team needs by reducing platform friction and improving self-service capabilities.

Requirements

  • Experience designing, building, and running distributed cloud architectures and large-scale web-based production systems.
  • Deep knowledge of Kubernetes, including operating production-grade clusters, writing custom controllers or operators, and tuning autoscaling.
  • Experience with AWS and GCP.
  • Baseline coding ability, with primary exposure to Python and some C++.
  • Comfort operating close to production, including debugging incidents and improving reliability and observability.
  • Strong systems thinking with attention to edge cases, failure modes, and implementation details.
  • Focus on performance, reliability, cost efficiency, and operational simplicity.
  • Ownership mindset and willingness to tackle broken systems and follow through.
  • SQL experience and curiosity about real-time analytical systems; ClickHouse experience or launching database systems at scale is a strong plus.
  • Familiarity with Traefik, Varnish, Redis, Terraform, or Ansible is helpful.
  • Clear written communication for async work and documentation.
  • Use of AI tools such as Claude Code, Cursor, and ChatGPT to improve workflows.
  • Fluent in English and Spanish.
  • Willingness to participate in on-call rotations.
  • Located in an EU timezone.

Benefits

  • Remote-first work environment.
  • Occasional in-person meetups, especially in Madrid.
  • Influence over how Tinybird operates and scales.
  • Opportunity to work closely with product, support, customer success, and engineering teams.
  • Recruitment process designed to be simple and avoid unnecessary steps.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

Counterpart Health 51-200 hospital & health care

Counterpart Health is hiring a Senior Site Reliability and Infrastructure Engineer to support and evolve the technology platform behind its primary care tool and maintain reliable infrastructure for domestic and international workloads.

AWS Azure CI/CD Containerd DNS Docker GCP Go gRPC Helm Kubernetes Linux Load Balancing Prometheus Python Shell Scripting TCP/IP
15 hours, 28 minutes ago

Senior Test Platform & Reliability Engineer - Star Trek Fleet Command

Scopely 1K-5K Internet Software & Services

Scopely is hiring a Senior Test Platform & Reliability Engineer in Ireland to build validation, reliability, and developer enablement platforms for Star Trek Fleet Command’s large-scale live-service backend systems.

AWS Bash CI/CD Docker GitLab Go Python Terraform
15 hours, 43 minutes ago

Senior Software Engineer - Databases, SRE | Canada | Remote

Grafana 1K-5K IT Services

Grafana Labs is hiring a Senior Software Engineer for its remote SRE team to improve reliability and operability of Grafana Cloud database services for high-SLA customers across AWS, GCP, and Azure.

AWS Azure GCP Go Helm Java Kubernetes Linux Microservices Python Terraform
1 day, 14 hours ago

Senior Site Reliability Engineer

Semios 51-250 Food Products

Semios Group is hiring a Senior Site Reliability Engineer to help scale, secure, and improve the reliability of its global agricultural technology platform.

AWS Azure Bash Buildkite CI/CD Datadog Docker Envoy GCP Git GitHub GitHub Actions GitLab Go Jenkins Kubernetes Linux NATS New Relic Prometheus Python Ruby Splunk Terraform
1 day, 15 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers