Boundless

Boundless

Boundless provides a protocol for accessing verifiable compute across various blockchain networks, enabling rapid upgrades to zero-knowledge rollups while ensuring transparency and security in trade settlements.

Internet Software & Services
Founded 2005

Description

  • Operate a heterogeneous, multi-region GPU fleet across consumer and datacenter hardware.
  • Build reliable scheduling patterns for inference workloads across cloud and on-prem infrastructure.
  • Maximize GPU utilization by managing workload placement across spot, on-prem, and cloud capacity.
  • Tune bare-metal and GPU performance to improve throughput per node.
  • Build secure fleet access and operational tooling using systems like Tailscale and Teleport.
  • Implement robust observability, alerting, and zero-downtime rollouts across the fleet.
  • Drive cost reduction by lowering GPU-hour spend without sacrificing reliability.
  • Work autonomously to navigate ambiguity and improve infrastructure systems proactively.

Requirements

  • 5+ years of infrastructure or DevOps experience operating large-scale production systems.
  • Deep expertise in Kubernetes, Docker, and container orchestration at scale.
  • Strong Linux systems administration skills.
  • Proficiency with infrastructure-as-code tools such as Terraform, Ansible, or Pulumi.
  • Track record of managing mission-critical, high-throughput systems.
  • Strong infrastructure-as-code background in heterogeneous environments.
  • Proficiency in at least one scripting or programming language such as Python, Bash, TypeScript, or Go.
  • Comfort navigating ambiguity with a strong bias for action.
  • Experience with GPU computing infrastructure, including CUDA, bare-metal optimization, or kernel tuning (preferred).
  • Experience operating ML training or other large-scale distributed compute infrastructure (preferred).
  • Experience with GPU fleet orchestration tools such as SkyPilot, Ray, or Slurm (preferred).
  • Familiarity with fleet access and networking tools such as Tailscale or Teleport (preferred).
  • Knowledge of network optimization and topology design (preferred).
  • Experience with multi-region, globally distributed systems (preferred).
  • Proficiency in Rust or low-level systems programming (preferred).
  • Experience with on-premises data center operations (preferred).
  • Candidates must include a public GitHub profile with at least 1 year of activity/history.

Benefits

  • Competitive salary of approximately US$175k to $250k annually.
  • Equity allocation.
  • Health, dental, and vision coverage for U.S. employees, with region-adjusted coverage globally.
  • Flexible PTO.
  • Professional development and conference travel budget.
  • Remote-first work environment.
  • Regular off-sites and a high-trust, high-velocity team environment.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Cloud Security Engineer

Smile Digital Health 251-1K IT Services

Smile Digital Health is seeking a Cloud Security Engineer to design, automate, deploy, and support secure production-grade infrastructure for healthcare data platforms across AWS, Azure, OCI, and GCP.

Ansible AWS Azure Docker HIPAA Kubernetes OpenShift Terraform
1 day, 21 hours ago

Director of Infrastructure

Sysdig 251-1K IT Services

Sysdig is hiring an infrastructure leader to own global cloud strategy, reliability, cost efficiency, and team execution for its large-scale, multi-tenant cloud security platform.

AWS Azure Kubernetes Pulumi Terraform
1 day, 21 hours ago

Cloud Infrastructure Engineer (Remote LATAM)

Atmosera 51-250 IT Services

Atmosera is seeking a senior Azure Cloud Infrastructure Engineer contractor to design, implement, migrate, and operate secure cloud environments while advising clients and supporting internal operations.

Agile Ansible Azure PowerShell SQL Server Terraform Windows Server
3 days, 21 hours ago

Managed Services Operations Lead - FinOps

AHEAD 1K-5K IT Services

AHEAD is seeking a remote Managed Services Operations Lead, FinOps to serve as the senior technical authority for its Cloud FinOps practice, leading enterprise solutions, client escalations, delivery standards, and AI-enabled service development across AWS, Azure, and GCP.

AWS Azure Bash Databricks GCP JIRA Kubernetes Machine Learning Power BI PowerShell Python Sentinel Serverless Snowflake SQL Tableau Terraform
3 days, 22 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers