Simulmedia

Simulmedia

Simulmedia is a pioneer in cross-channel TV advertising, offering advertisers the reach of linear TV combined with digital targeting. Their VAMOS platform enables intelligent audience optimization and direct sales attribution, revolutionizing TV advert...

Media
51-250
Founded 2008
$87M raised

Description

  • Design and build batch data pipelines that ingest, validate, and transform multi-billion-row datasets from external providers and internal systems.
  • Model complex data structures, including dimensional, reference, and temporal models, and evolve them safely as upstream schemas change.
  • Develop and operate workloads on Databricks/Spark/Delta and Redshift, including migrating pipelines to the lakehouse platform.
  • Orchestrate pipelines in Airflow, including scheduling, dependencies, retries, backfills, and alerting.
  • Design parity checks and reconciliation queries, run historical backfills, and investigate data discrepancies at the row level.
  • Build and maintain Python services and REST APIs that serve data to internal products.
  • Optimize performance and cost through query tuning, table design, workload management, and compute right-sizing.
  • Monitor production pipelines, participate in incident triage and root-cause analysis, and harden systems against repeat failures.
  • Collaborate with product managers, data scientists, and other stakeholders to deliver roadmap items.
  • Work in an Agile team and experiment with new technologies to improve software and data workflows.

Requirements

  • Bachelor's degree in Computer Science, Computer Engineering, a relevant technical field, or equivalent practical experience.
  • 7+ years of work experience as a data engineer.
  • Proficiency in Python as the primary development language in recent years.
  • Expert-level SQL for complex analytical queries, performance tuning, and debugging result discrepancies.
  • Hands-on experience with a distributed data processing platform such as Spark/Databricks, EMR, Snowflake, or BigQuery.
  • Experience with a columnar data warehouse such as Redshift, Snowflake, BigQuery, or ClickHouse.
  • Ability to design complex data models, including normalized, dimensional, and temporal models.
  • Experience with workflow orchestration tools such as Airflow, including DAGs, dependencies, and backfills.
  • Experience integrating third-party data feeds, including schema drift, late deliveries, missing data, and vendor data-quality defects.
  • Experience building REST services in Python using FastAPI, Flask, or similar frameworks.
  • Working knowledge of AWS services such as S3, IAM, and ECS, plus Docker.
  • Good knowledge of engineering best practices, including unit testing, integration testing, code review, and CI/CD.
  • Must be able to communicate with U.S.-based teams and work 11:00 AM — 8:00 PM EEST.
  • Experience with Delta Lake or medallion lakehouse architectures is a plus.
  • Experience migrating legacy pipelines between platforms with strict parity requirements is a plus.
  • Experience with advertising, media, or measurement industry data is a plus.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Lead Data Engineer

Media.Monks 5K-10K Media

Monks Technology Services is seeking a fully remote Lead Data Engineer for a fixed-term contract to lead the modernization of legacy Python/Spark ETL pipelines into dbt Core models on Databricks supporting analytics and business intelligence.

Apache Spark CI/CD Databricks Git GitHub Actions Python SQL
17 hours, 34 minutes ago

Staff Data Engineer

Webflow 251-1K Internet Software & Services

Webflow is hiring a Staff Data Engineer to lead data platform initiatives involving its data lake, event instrumentation, data quality, governance, and cloud infrastructure for modern web marketing operations.

Apache Airflow Apache Spark CI/CD Kafka SQL
17 hours, 34 minutes ago

Data Engineer [Zeal]

Livefront 11-50 Internet Software & Services

Zeal, now part of Livefront, is hiring a Data Engineer across its U.S. hubs to build reliable data pipelines and architectures that support technology consulting engagements for Fortune 1000 companies.

Apache Airflow CI/CD Databricks dbt GCP Git Kafka Oracle PostgreSQL Power BI Python RabbitMQ Snowflake SQL SQL Server Tableau
17 hours, 49 minutes ago

Sr Data Ops Engineer

Coderio 51-250 Internet Software & Services

Coderio is hiring a DataOps Engineer and Technical Referent to work with international customers on designing, implementing, and operating scalable cloud data infrastructure and automation solutions.

AWS Bash dbt Docker GitHub Python SQL Terraform
18 hours, 19 minutes ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers