Capgemini

Capgemini

Capgemini is a global leader in consulting, technology services, and digital transformation, empowering businesses with innovative solutions and expertise to thrive in a rapidly evolving market.

Internet Software & Services
100K+
Founded 1967
$93M raised

Description

  • Design, build, and manage end-to-end data pipelines across the bronze, silver, and gold layers of the medallion architecture.
  • Ingest and process raw data using Spark and Amazon EMR for scalable distributed computation.
  • Develop and automate data transformations for the base vault using DBT.
  • Support business vault and gold-layer data modeling to produce curated datasets.
  • Work with orchestration tools such as Airflow or AWS Step Functions to run and manage data workflows.
  • Use Amazon S3 and Apache Iceberg to store and manage data assets.
  • Collaborate on data architecture practices that support governance, quality, and reusable data products.
  • Maintain and optimize Elasticsearch-related data engineering components, including indexing, query performance, and administration.

Requirements

  • At least 5 years of experience as an Elasticsearch Data Engineer with ELK stack expertise.
  • Strong experience with Elasticsearch cluster optimization, query development, data modeling, performance tuning, and administration (4–6 years).
  • Deep experience with Spark, Python, ETLs, and Amazon EMR.
  • Hands-on experience with DBT for data transformation and modeling.
  • Experience with Apache Airflow, AWS Step Functions, or similar orchestration tools.
  • Expert knowledge of Amazon S3 and Apache Iceberg for data storage and management.
  • Experience with Kubernetes for container orchestration.
  • Experience with Dremio, Looker, or equivalent semantic layer/business view technologies.
  • Intermediate AWS Cloud experience, including AWS Lambda, Step Functions, IAM, SNS, API Gateway, VPC, and Transit Gateway.
  • BS in Computer Science, Data Engineering, Data Modeling, or a similar field; Big Data or AWS certification is a plus.
  • Full English fluency.
  • Strong analytical, problem-solving, and communication skills.
  • Java Spring Boot and IBM ACE programming experience.

Benefits

  • Competitive salary with performance-based bonuses.
  • Comprehensive benefits package.
  • Private health insurance.
  • Pension plan.
  • Paid time off.
  • Career development, training, and development opportunities.
  • Flexible work arrangements, including remote and/or office-based options.
  • Dynamic and inclusive work culture within a globally recognized group.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Lead Data Engineer

Media.Monks 5K-10K Media

Monks Technology Services is seeking a fully remote Lead Data Engineer for a fixed-term contract to lead the modernization of legacy Python/Spark ETL pipelines into dbt Core models on Databricks supporting analytics and business intelligence.

Apache Spark CI/CD Databricks Git GitHub Actions Python SQL
18 hours, 1 minute ago

Staff Data Engineer

Webflow 251-1K Internet Software & Services

Webflow is hiring a Staff Data Engineer to lead data platform initiatives involving its data lake, event instrumentation, data quality, governance, and cloud infrastructure for modern web marketing operations.

Apache Airflow Apache Spark CI/CD Kafka SQL
18 hours, 1 minute ago

Data Engineer [Zeal]

Livefront 11-50 Internet Software & Services

Zeal, now part of Livefront, is hiring a Data Engineer across its U.S. hubs to build reliable data pipelines and architectures that support technology consulting engagements for Fortune 1000 companies.

Apache Airflow CI/CD Databricks dbt GCP Git Kafka Oracle PostgreSQL Power BI Python RabbitMQ Snowflake SQL SQL Server Tableau
18 hours, 17 minutes ago

Sr Data Ops Engineer

Coderio 51-250 Internet Software & Services

Coderio is hiring a DataOps Engineer and Technical Referent to work with international customers on designing, implementing, and operating scalable cloud data infrastructure and automation solutions.

AWS Bash dbt Docker GitHub Python SQL Terraform
18 hours, 47 minutes ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers