NTT DATA

NTT DATA

NTT DATA GROUP CORPORATION is a global IT and business services provider headquartered in Tokyo, Japan. It offers a range of services including cloud, cybersecurity, data and intelligence, consulting, and managed services, and is managed under the NTT ...

IT Services
100K+
Founded 1988

Description

  • Design, build, and maintain ingestion pipelines from relational, file, API, and streaming sources into Amazon S3.
  • Define and maintain landing and raw zone layouts, including partitioning, file formats, naming conventions, compression, retention, and source-data immutability.
  • Onboard new source systems end to end, including connectivity, network configuration, credentials, extraction strategy, and coordination with source owners.
  • Ensure load completeness and integrity through reconciliation, control tables, reprocessing procedures, and handling of failed or partial runs.
  • Implement processing across raw, cleansed, and gold layers of the medallion architecture.
  • Register and maintain datasets in the Glue Data Catalog and keep schemas, partitions, and metadata aligned across AWS services.
  • Tune job performance and cost by monitoring execution time, DPU consumption, and file layout.
  • Embed data quality controls, define rulesets, handle rejected records, and publish quality metrics.
  • Orchestrate multi-step pipelines using Glue Workflows, Step Functions, and EventBridge.
  • Provide L2/L3 production support, including incident analysis, root cause investigation, backfills, and ongoing improvements.

Requirements

  • 5+ years of data engineering experience.
  • Working proficiency in English and Spanish, both written and spoken.
  • Hands-on experience building ETL pipelines with AWS Glue and PySpark.
  • Experience moving data from databases and files into Amazon S3.
  • Strong knowledge of Python, SQL, Spark, and data transformation.
  • Experience designing reliable, scalable, and restartable data pipelines.
  • Understanding of S3 data lakes and layered data architectures.
  • Experience with AWS workflow, monitoring, and data-quality tools.
  • Experience supporting production pipelines and troubleshooting failures.
  • Experience with AWS data services such as DMS, Athena, Redshift, or Lake Formation is preferred.
  • Familiarity with Informatica, Denodo, Terraform, or CloudFormation is preferred.
  • Experience in enterprise or highly regulated environments is preferred.
  • Relevant AWS, data engineering, or database certification is preferred.

Benefits

  • Remote opportunity for LATAM-based candidates.
  • Opportunity to work with a global U.S. client.
  • Career development model focused on empowerment and rewards.
  • Professional growth in a young, fast-growing, innovative company.
  • Commitment to diversity and equal opportunity employment.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Data Engineer (6278)

Dan.com - a GoDaddy brand Internet Software & Services

itD is hiring a remote Data Engineer to build and maintain scalable data infrastructure for AI product analytics and topline metrics across product surfaces.

AWS Azure GCP Generative AI Machine Learning Python Snowflake SQL
29 minutes ago

Lead Data Manager/SAS programmer (Job 1425)

DLH 1K-5K Construction & Engineering

DLH is seeking a Lead Data Manager/SAS Programmer to lead data management and statistical programming for federally funded and commercially funded clinical research and epidemiological studies.

GCP Git HIPAA Python R SQL
29 minutes ago

Data Engineer with Airflow

Xebia 1K-5K Internet Software & Services

Xebia is hiring a Data Engineer to support a managed Apache Airflow platform, helping customer data teams adopt, integrate, and operate Airflow in production environments.

Apache Airflow AWS Azure Databricks dbt GCP Kubernetes Python Snowflake SQL Terraform
29 minutes ago

Staff Software Engineer- Data Ingestion

Aledade 1K-5K Health Care Providers & Services

Staff Software Engineer at a data-focused company, leading backend architecture modernization and the evolution of distributed systems that power scalable data storage, processing, and analytics.

Apache Spark AWS Azure C# C++ CI/CD Databricks Docker GCP Go Java Kubernetes Python Scala Snowflake SQL
1 hour, 14 minutes ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers