Sr. Staff Platform/Data Reliability Engineer, Databricks (R5537)

3 weeks, 4 days ago
Full-time
Lead
DevOps and Infrastructure
Bitly

Bitly

Bitly is a link management platform offering URL shortening, QR codes, and a Link in Bio solution for brands to optimize customer experience.

Internet Software & Services
51-250
Founded 2008
$92M raised

Description

  • Own operational excellence for the Databricks platform, including monitoring, alerting, observability, incident response support, and runbooks.
  • Define and maintain CI/CD and promotion standards for Databricks assets across dev-to-prod environments.
  • Design platform standards for job orchestration, cluster and compute policies, service principals, and execution reliability.
  • Create reusable operational templates and onboarding patterns for new Databricks domains.
  • Partner with data engineering to ensure ingestion and medallion patterns are observable, recoverable, cost-aware, and secure in production.
  • Work with cloud and infrastructure teams to align Databricks usage with enterprise cloud standards.
  • Help enforce technical controls for data segregation, access boundaries, and operational compliance.
  • Track and improve platform health metrics such as job success rates, incidents, reliability, cost efficiency, and drift.
  • Document platform standards, operational expectations, and support models.
  • Mentor engineers growing into platform responsibilities.

Requirements

  • 12+ years of relevant experience in data platform engineering, platform operations, site reliability engineering, or cloud data infrastructure.
  • Hands-on experience with Databricks or a closely related cloud data platform in production.
  • Experience designing or operating CI/CD, environment promotion, version control, and deployment automation for data platforms and pipelines.
  • Strong understanding of observability, monitoring, alerting, incident management, and reliability engineering.
  • Experience with compute policy design, workload isolation, service principals, and secure production execution patterns.
  • Ability to work in a regulated or security-sensitive environment with strong access control and auditability requirements.
  • Strong collaboration skills with cloud/infrastructure, security, data engineering, and analytics stakeholders.
  • Databricks certification and/or strong expertise with Delta Lake, Unity Catalog, Workflows, and Databricks Asset Bundles (preferred).
  • Experience with infrastructure-as-code and platform automation in enterprise environments (preferred).
  • Experience supporting commercial and government or otherwise segregated environments with different compliance and access requirements (preferred).
  • Experience in defense, aerospace, federal, or another regulated industry (preferred).

Benefits

  • Pay within range listed plus bonus, benefits, and equity for full-time regular employees.
  • Temporary employee package includes pay within range plus a temporary benefits package after 60 days of employment.
  • Salary offers are based on experience, licenses/certifications, and work location.
  • All offers are contingent on a cleared background check and possible reference check.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

SRE [Antifraud]

Banco Plata, S.A., Institución de Banca Múltiple. 1001-5000 Banking / financial services

Join the Antifraud team as an SRE, helping operate scalable AML, Anti-Fraud, and KYC platforms that perform real-time financial crime prevention checks without slowing transactions.

Bash ELK Stack Helm Kafka Kubernetes Load Balancing Prometheus Python Secrets Management
10 hours, 16 minutes ago

Staff Site Reliability Engineer

Zscaler 1K-5K Internet Software & Services

Zscaler is seeking a remote Staff Site Reliability Engineer in the Netherlands to build, secure, automate, and operate scalable Linux, Kubernetes, and cloud infrastructure for its global security platform.

Ansible Bash DHCP Docker Go HashiCorp Vault Kubernetes Linux Python Secrets Management SSH
1 day, 11 hours ago

Staff Platform Engineer

Catena Clearing Internet Software & Services

Catena is hiring a Staff-level infrastructure engineering leader to own the cloud infrastructure, distributed event pipelines, and developer tooling behind its real-time universal fleet telematics data platform.

AWS AWS CDK Docker Kafka Microservices OAuth PostgreSQL Pulumi Python Redis REST API Terraform
1 day, 11 hours ago

Staff Site Reliability Engineer-Observability

GoDaddy 5001-10000 Technology, Information and Internet

GoDaddy is hiring a remote Reliability Engineer in India to operate and modernize the monitoring, compliance, infrastructure, and incident-response systems supporting its global Domains platform.

Ansible Argo CD AWS AWS CDK CI/CD Cybersecurity Elasticsearch Git Go Gradle Grafana Jenkins Kubernetes Maven MySQL PostgreSQL Prometheus Python Ruby SQL Terraform
2 days, 11 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers