Capgemini

Capgemini

Capgemini is a global leader in consulting, technology services, and digital transformation, empowering businesses with innovative solutions and expertise to thrive in a rapidly evolving market.

Internet Software & Services
100K+
Founded 1967
$93M raised

Description

  • Design, build, and manage end-to-end data pipelines across the bronze, silver, and gold layers of the medallion architecture.
  • Ingest and process raw data using Spark and Amazon EMR for scalable distributed computation.
  • Develop and automate data transformations for the base vault using DBT.
  • Support business vault and gold-layer data modeling to produce curated datasets.
  • Work with orchestration tools such as Airflow or AWS Step Functions to run and manage data workflows.
  • Use Amazon S3 and Apache Iceberg to store and manage data assets.
  • Collaborate on data architecture practices that support governance, quality, and reusable data products.
  • Maintain and optimize Elasticsearch-related data engineering components, including indexing, query performance, and administration.

Requirements

  • At least 5 years of experience as an Elasticsearch Data Engineer with ELK stack expertise.
  • Strong experience with Elasticsearch cluster optimization, query development, data modeling, performance tuning, and administration (4–6 years).
  • Deep experience with Spark, Python, ETLs, and Amazon EMR.
  • Hands-on experience with DBT for data transformation and modeling.
  • Experience with Apache Airflow, AWS Step Functions, or similar orchestration tools.
  • Expert knowledge of Amazon S3 and Apache Iceberg for data storage and management.
  • Experience with Kubernetes for container orchestration.
  • Experience with Dremio, Looker, or equivalent semantic layer/business view technologies.
  • Intermediate AWS Cloud experience, including AWS Lambda, Step Functions, IAM, SNS, API Gateway, VPC, and Transit Gateway.
  • BS in Computer Science, Data Engineering, Data Modeling, or a similar field; Big Data or AWS certification is a plus.
  • Full English fluency.
  • Strong analytical, problem-solving, and communication skills.
  • Java Spring Boot and IBM ACE programming experience.

Benefits

  • Competitive salary with performance-based bonuses.
  • Comprehensive benefits package.
  • Private health insurance.
  • Pension plan.
  • Paid time off.
  • Career development, training, and development opportunities.
  • Flexible work arrangements, including remote and/or office-based options.
  • Dynamic and inclusive work culture within a globally recognized group.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Data Engineer (Contract)

Concurrency 51-250 Internet Software & Services

Concurrency is hiring a Data Engineer for its Cloud, Data & AI Solutions Team to deliver end-to-end Microsoft Azure and Fabric data solutions that support client digital transformation.

Apache Spark Azure CI/CD Git Machine Learning Power BI Python SQL
2 hours, 3 minutes ago

Data Engineer II

Netomi 51-250 IT Services

Netomi is hiring a Data Engineer in Gurugram to build reliable data infrastructure and analytics solutions that improve AI-powered enterprise customer experiences and product operations.

Apache Airflow Apache Spark CI/CD Databricks Docker HIPAA Kafka Kubernetes Luigi MySQL PostgreSQL Prefect Python RabbitMQ Snowflake SQL
2 hours, 3 minutes ago

[Job-31682] Senior Data Engineer [Databricks]

CI&T 5K-10K Internet Software & Services

A CI&T busca uma pessoa engenheira de dados para estruturar e acelerar a plataforma de dados de um grande varejista farmacêutico brasileiro, modernizando a orquestração, a qualidade dos dados e a distribuição de produtos baseados em Databricks.

Apache Airflow Apache Spark AWS Databricks SFTP Terraform
2 hours, 18 minutes ago

Data Engineer

payabl. 51-250 Diversified Financial Services

As a Data Engineer at payabl., you will build and maintain reliable data pipelines and business-ready datasets that support analytics and decision-making across the company’s global payments platform.

Apache Airflow Apache Spark AWS ClickHouse Dagster Databricks dbt Docker Git Kafka Kubernetes MariaDB MongoDB MySQL PostgreSQL Power BI Python Snowflake SQL Tableau Terraform
2 hours, 18 minutes ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers