Sonatype

Sonatype

Sonatype provides secure software development solutions by leveraging open source and artificial intelligence, ensuring that organizations can build applications quickly and safely through automated governance, policy enforcement, and comprehensive mon...

Internet Software & Services
51-250
Founded 2008
$155M raised

Description

  • Design, build, and maintain scalable data pipelines and ETL/ELT processes.
  • Architect and optimize data models and storage solutions for analytics and operational use.
  • Collaborate with data scientists, analysts, engineers, and other stakeholders to deliver trusted datasets.
  • Own and evolve parts of the data platform using Databricks and Spark.
  • Implement observability, alerting, and data quality monitoring for critical pipelines.
  • Drive engineering best practices in documentation, testing, and CI/CD.
  • Help define long-term data platform architecture and mentor the team.
  • Contribute to the design and evolution of the next-generation data lakehouse architecture.

Requirements

  • 8+ years of experience as a Data Engineer or in a similar backend engineering role.
  • Bachelor’s degree in Computer Science, Engineering, or a related technical field.
  • Strong programming skills in Python, Scala, or Java.
  • Hands-on experience with distributed data systems like Spark or Kafka.
  • Proficiency in complex SQL and NoSQL query development and performance optimization.
  • Experience building and maintaining robust ETL/ELT pipelines in production.
  • Understanding of data modeling techniques such as star schema and dimensional modeling.
  • Experience leveraging AI-assisted development tools and AI/ML technologies to improve data engineering workflows, developer productivity, data quality, and operations.
  • Experience tuning Spark jobs, optimizing joins, and managing Delta Lake architecture for batch and streaming data.
  • Preferred: Familiarity with software supply chain, cybersecurity, or large-scale software ecosystem data.
  • Preferred: Track record of improving data platform reliability, scalability, performance, and cost efficiency.
  • Preferred: Familiarity with workflow orchestration tools such as Airflow or Dagster.
  • Preferred: Hands-on experience with cloud data platforms, particularly AWS.
  • Preferred: Familiarity with modern table formats such as Delta Lake, Apache Iceberg, or Apache Hudi.
  • Preferred: Experience implementing data observability, lineage, governance, and automated data quality frameworks.
  • Preferred: Experience designing real-time or streaming data architectures using data lake technologies.

Benefits

  • Flexible working practices and remote-friendly flexibility.
  • Parental leave policy.
  • Diversity and inclusion working groups.
  • Paid Volunteer Time Off (VTO).
  • Opportunity to work on high-impact problems in software supply chain security.
  • Work with modern open-source and cloud-native technologies.
  • Collaborative culture that values learning, autonomy, and impact.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Lead Data Engineer

Media.Monks 5K-10K Media

Monks Technology Services is seeking a fully remote Lead Data Engineer for a fixed-term contract to lead the modernization of legacy Python/Spark ETL pipelines into dbt Core models on Databricks supporting analytics and business intelligence.

Apache Spark CI/CD Databricks Git GitHub Actions Python SQL
17 hours, 8 minutes ago

Staff Data Engineer

Webflow 251-1K Internet Software & Services

Webflow is hiring a Staff Data Engineer to lead data platform initiatives involving its data lake, event instrumentation, data quality, governance, and cloud infrastructure for modern web marketing operations.

Apache Airflow Apache Spark CI/CD Kafka SQL
17 hours, 8 minutes ago

Data Engineer [Zeal]

Livefront 11-50 Internet Software & Services

Zeal, now part of Livefront, is hiring a Data Engineer across its U.S. hubs to build reliable data pipelines and architectures that support technology consulting engagements for Fortune 1000 companies.

Apache Airflow CI/CD Databricks dbt GCP Git Kafka Oracle PostgreSQL Power BI Python RabbitMQ Snowflake SQL SQL Server Tableau
17 hours, 23 minutes ago

Sr Data Ops Engineer

Coderio 51-250 Internet Software & Services

Coderio is hiring a DataOps Engineer and Technical Referent to work with international customers on designing, implementing, and operating scalable cloud data infrastructure and automation solutions.

AWS Bash dbt Docker GitHub Python SQL Terraform
17 hours, 53 minutes ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers