Sonatype

Sonatype

Sonatype provides secure software development solutions by leveraging open source and artificial intelligence, ensuring that organizations can build applications quickly and safely through automated governance, policy enforcement, and comprehensive mon...

Internet Software & Services
51-250
Founded 2008
$155M raised

Description

  • Design, build, and maintain scalable data pipelines and ETL/ELT processes.
  • Architect and optimize data models and storage solutions for analytics and operational use.
  • Collaborate with data scientists, analysts, engineers, and other stakeholders to deliver trusted datasets.
  • Own and evolve parts of the data platform using Databricks and Spark.
  • Implement observability, alerting, and data quality monitoring for critical pipelines.
  • Drive engineering best practices in documentation, testing, and CI/CD.
  • Help define long-term data platform architecture and mentor the team.
  • Contribute to the design and evolution of the next-generation data lakehouse architecture.

Requirements

  • 8+ years of experience as a Data Engineer or in a similar backend engineering role.
  • Bachelor’s degree in Computer Science, Engineering, or a related technical field.
  • Strong programming skills in Python, Scala, or Java.
  • Hands-on experience with distributed data systems like Spark or Kafka.
  • Proficiency in complex SQL and NoSQL query development and performance optimization.
  • Experience building and maintaining robust ETL/ELT pipelines in production.
  • Understanding of data modeling techniques such as star schema and dimensional modeling.
  • Experience leveraging AI-assisted development tools and AI/ML technologies to improve data engineering workflows, developer productivity, data quality, and operations.
  • Experience tuning Spark jobs, optimizing joins, and managing Delta Lake architecture for batch and streaming data.
  • Preferred: Familiarity with software supply chain, cybersecurity, or large-scale software ecosystem data.
  • Preferred: Track record of improving data platform reliability, scalability, performance, and cost efficiency.
  • Preferred: Familiarity with workflow orchestration tools such as Airflow or Dagster.
  • Preferred: Hands-on experience with cloud data platforms, particularly AWS.
  • Preferred: Familiarity with modern table formats such as Delta Lake, Apache Iceberg, or Apache Hudi.
  • Preferred: Experience implementing data observability, lineage, governance, and automated data quality frameworks.
  • Preferred: Experience designing real-time or streaming data architectures using data lake technologies.

Benefits

  • Flexible working practices and remote-friendly flexibility.
  • Parental leave policy.
  • Diversity and inclusion working groups.
  • Paid Volunteer Time Off (VTO).
  • Opportunity to work on high-impact problems in software supply chain security.
  • Work with modern open-source and cloud-native technologies.
  • Collaborative culture that values learning, autonomy, and impact.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

[Job-31682] Senior Data Engineer [Databricks]

CI&T 5K-10K Internet Software & Services

A CI&T busca uma pessoa engenheira de dados para estruturar e acelerar a plataforma de dados de um grande varejista farmacêutico brasileiro, modernizando a orquestração, a qualidade dos dados e a distribuição de produtos baseados em Databricks.

Apache Airflow Apache Spark AWS Databricks SFTP Terraform
12 minutes ago

Data Engineer

payabl. 51-250 Diversified Financial Services

As a Data Engineer at payabl., you will build and maintain reliable data pipelines and business-ready datasets that support analytics and decision-making across the company’s global payments platform.

Apache Airflow Apache Spark AWS ClickHouse Dagster Databricks dbt Docker Git Kafka Kubernetes MariaDB MongoDB MySQL PostgreSQL Power BI Python Snowflake SQL Tableau Terraform
12 minutes ago

Product Engineer — Monetization & Commerce Data

Terrific 1-10 Internet Software & Services

Terrific is hiring a remote Product Engineer in Europe to define and build its monetization and commerce-data platform for retailer and publisher advertising, turning first-party shopper interactions into measurable revenue.

E-commerce GCP Node.js SQL TypeScript
12 minutes ago

[Job-31419] Mid-level Data Developer, Colombia

CI&T 5K-10K Internet Software & Services

CI&T is seeking a Data Developer to build and maintain ingestion pipelines into OneLake for ATL Technology’s Foundation Pod across multiple enterprise source systems, supporting the Data Engineer Lead.

Apache Spark CI/CD Databricks Python SQL
12 minutes ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers