Data Engineer Sr. Consultant

NTT DATA

3–6 yrs Bengaluru Full Time Hybrid (office + remote)
NTT DATA logo
Posted : today
Actively hiring

Job description

Join NTT DATA's innovative team as a Data Engineer Sr. Consultant in Bangalore, India. We seek passionate individuals to drive advancements in data engineering and contribute to impactful client success.

This role involves transforming existing ETL workloads to Databricks Lakehouse, developing robust data pipelines, and implementing advanced data processing layers. You will work with cutting-edge technologies to migrate and optimize data solutions, ensuring efficiency and scalability.

As part of a forward-thinking organization, you'll collaborate with senior developers and technical leads, adhering to high standards of coding, performance, and security. We encourage a culture of continuous learning and collaboration through code reviews and knowledge sharing.

Responsibilities

Analyze and understand current Informatica workflows and AWS Glue jobs, focusing on source-to-target mappings and transformation logic. Develop Databricks notebooks and data pipelines to migrate existing ETL workloads to the Databricks Lakehouse environment. Implement Bronze, Silver, and Gold layer processing, adhering to established project architectures and coding patterns. Build batch and incremental processing pipelines utilizing Delta Lake and effective watermark/control mechanisms. Develop and maintain Databricks Jobs/Workflows and Lakeflow Declarative Pipelines as required, ensuring seamless execution and integration. Reuse common project frameworks for logging, auditing, data quality, error handling, and notifications to enhance pipeline robustness.

Convert Informatica transformations, mappings, lookups, and business rules into equivalent Databricks implementations, preserving all original business logic. Support the migration of AWS Glue jobs and database-based processing into Databricks. Conduct unit testing and perform source-to-target data reconciliation for all migrated pipelines. Investigate and resolve data, SQL, PySpark, and pipeline execution issues promptly. Support integration testing, regression testing, UAT, and production validation phases. Collaborate with senior developers and technical leads to address technical issues and dependencies effectively. Follow established coding, performance, security, and development standards diligently. Participate actively in code reviews and incorporate feedback to enhance code quality. Prepare comprehensive technical documentation and support knowledge transfer for migrated workloads. Assist with deployment, monitoring, production stabilization, and cutover activities to ensure successful project delivery.

Qualifications

A Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent practical experience is required. Candidates should possess 3–6 years of hands-on experience in data engineering, ETL development, or application/data platform development. Proven expertise with Databricks and Apache Spark is essential, alongside strong development skills in Python/PySpark and SQL.

Familiarity with Databricks notebooks, Spark SQL/Databricks SQL, Delta Lake, and Databricks Jobs/Workflows is expected. Experience developing batch or incremental ETL/ELT pipelines is necessary. Proficiency with at least one ETL technology such as Informatica PowerCenter, AWS Glue, or SSIS, and experience with relational databases and SQL-based data processing are key. A solid understanding of data transformation, source-to-target mapping, and ETL development principles is crucial.

Experience with unit testing, data validation, and troubleshooting is required. Basic knowledge of Bronze/Silver/Gold or other layered data architectures is beneficial. The ability to work within an established development framework and follow coding standards is important. Excellent analytical, problem-solving, and communication skills are a must.

Essential Skills

DatabricksApache SparkPythonPySparkSQLDatabricks notebooksSpark SQLDatabricks SQLDelta LakeDatabricks Jobs/WorkflowsETLELTInformatica PowerCenterAWS GlueSSISrelational databasesdata transformationsource-to-target mappingunit testingdata validationtroubleshootingBronze/Silver/Gold architecturecoding standardsanalytical skillsproblem-solvingcommunication skills

Good to Have

Databricks certificationLakeflow Declarative PipelinesDelta Live TablesUnity CatalogDatabricks Asset BundlesCI/CDAWS S3AWS GlueAWS RedshiftAuto Loaderstreaming ingestionCDCSCD Type 1SCD Type 2data quality frameworksaudit frameworkserror-handling frameworksInformatica PowerCenter mappingsInformatica PowerCenter workflowsAWS LambdaSpark performance optimizationDelta performance optimizationcloud data modernizationcloud data migration

Highlights

  • Actively hiring

More Details

RoleData Engineer Sr. Consultant
Employment TypeFull Time, Hybrid (office + remote)

About the Company

NTT DATA logo

NTT DATA

IT Consulting