Stack AV

Stack AV

Stack AV is a Pittsburgh-based autonomous trucking company focused on developing AI-powered self-driving technology for the freight and logistics industry. Founded in 2023 by a team with extensive experience in autonomous systems, Stack AV employs around 150 people and operates in 15 states. The company specializes in autonomous trucking solutions that address key challenges in transportation and supply chain management. Their offerings include self-driving truck technology that utilizes advanced AI, robotics, and machine learning, as well as supply chain optimization solutions aimed at enhancing efficiency and reliability in freight transportation. Stack AV prioritizes safety and operates with transparency and accountability, aiming to meet the critical needs of its customers in the transportation sector. Stack AV is supported by SoftBank Group Corp., which provides financial backing for its initiatives in autonomous trucking.

information technology & services
201-500
Founded 2023
$1000M raised

Description

  • Instrument systems that schedule and execute large-scale batch workloads across Kubernetes clusters.
  • Diagnose and triage job failures for internal customers.
  • Collaborate with teams across the company to understand workload requirements and improve platform capabilities.
  • Increase the reliability and velocity of systems and processes through automation.
  • Document operational actions and build runbooks as a knowledge base and foundation for automation.
  • Participate in an on-call rotation to uphold production service SLOs and SLAs.
  • Contribute to platform tooling, automation, and CI/CD workflows.

Requirements

  • Fundamental understanding of Linux operating system internals, TCP/IP networking, and storage subsystems.
  • Strong experience with Kubernetes and container orchestration in production-grade environments.
  • Ability to understand engineering design limitations and guide teams on scaling services within budget and performance goals.
  • Strong experience implementing and debugging cloud-native and open source tools such as Kubernetes, etcd, Prometheus, and OpenTelemetry.
  • Strong communication skills and the ability to work effectively in a diverse and distributed team.
  • Ability to work in a role that may be subject to U.S. national security, residence, citizenship, and export control requirements.
  • Experience supporting high-scale batch compute systems and workflow orchestration systems is preferred.
  • Experience working at the intersection of infrastructure, distributed systems, and developer experience is preferred.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Incident Commander

PENN Entertainment 10K-50K Hotels, Restaurants & Leisure

PENN Interactive is hiring an Incident Commander to join its site reliability team and lead incident response and service reliability efforts across its online and physical platforms.

Ansible AWS Docker Elasticsearch GCP Helm JIRA Kafka Kubernetes Linux MySQL PostgreSQL Prometheus Python Redis Terraform
54 minutes ago

Site Reliability Engineer

MyFitnessPal 10K-50K Health Care Providers & Services

MyFitnessPal is hiring a Software Engineer III, Site Reliability to own production reliability and delivery pipeline security for the PEAS team supporting automation, CI/CD, and self-service platforms.

AWS CI/CD Datadog GitHub Actions Go HIPAA Kubernetes Python Secrets Management Terraform TypeScript
1 hour, 39 minutes ago

Site Reliability Engineer 2 (Azure)

PhonePe 5K-10K Capital Markets

PhonePe Limited is hiring a Site Reliability Engineer to manage and scale core cloud infrastructure for a high-volume digital payments environment in India.

Ansible Azure Bash DNS Docker Go Grafana HAProxy InfluxDB Java Linux MySQL Nginx Prometheus Python RabbitMQ SaltStack Terraform Ubuntu
1 day ago

Sr. Control System Engineer/Site Reliability Engineer (SRE)

QuEra Computing 11-50 Internet Software & Services

QuEra is seeking a Sr. Control System Engineer/Site Reliability Engineer to integrate and maintain the hardware and software systems that support its quantum control stack and keep development and production environments reliable.

Ansible Bash CI/CD Debian DHCP DNS Docker ELK Stack Embedded Systems Git GitLab CI Go Grafana Jenkins Kubernetes Linux Prometheus Python TCP/IP Terraform Ubuntu
1 day, 1 hour ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers