Data Platform Modernization Engineer - Junior (Databricks)
NTT DATA
NTT DATA
Join NTT DATA and contribute to an enterprise data platform modernization initiative as a Junior Data Platform Modernization Engineer specializing in Databricks. This role focuses on migrating existing ETL workloads from legacy systems like Informatica PowerCenter and AWS Glue to the Databricks Lakehouse. You will develop scalable and reliable data pipelines, leveraging established migration patterns and frameworks.
This is an opportunity to work with cutting-edge technologies and play a key role in transforming data infrastructure. The ideal candidate possesses strong hands-on development skills in PySpark, Python, and SQL, with a solid background in ETL development, data transformation, testing, and troubleshooting.
Analyze existing Informatica and AWS Glue processes to understand data flows and transformation logic. Develop Databricks notebooks and pipelines for migrating ETL workloads to the Databricks Lakehouse. Implement data transformations using Python, PySpark, Spark SQL, and Databricks SQL. Build Bronze, Silver, and Gold data layers following project architecture and coding standards. Create batch and incremental processing pipelines using Delta Lake and effective control mechanisms. Develop and maintain Databricks Jobs/Workflows and Lakeflow Declarative Pipelines. Utilize common project frameworks for logging, auditing, data quality, and error handling. Translate Informatica transformations and business rules into Databricks implementations. Support the migration of AWS Glue jobs and database processing to Databricks. Perform unit testing and data reconciliation for migrated pipelines. Investigate and resolve data, SQL, PySpark, and pipeline execution issues. Assist with integration testing, UAT, and production validation. Collaborate with senior developers to resolve technical challenges. Adhere to coding, performance, and security standards. Participate in code reviews and incorporate feedback. Prepare technical documentation and support knowledge transfer. Assist with deployment, monitoring, and production stabilization.
A minimum of 3–6 years of experience in data engineering, ETL development, or application/data platform development is required. Proficiency with Databricks and Apache Spark is essential. Strong development skills in Python/PySpark and SQL are a must. Experience with Databricks notebooks, Spark SQL/Databricks SQL, Delta Lake, and Databricks Jobs/Workflows is expected. You should have a proven track record in developing batch or incremental ETL/ELT pipelines. Familiarity with ETL technologies like Informatica PowerCenter, AWS Glue, or SSIS is necessary. Experience working with relational databases and SQL-based data processing is crucial. A good understanding of data transformation, source-to-target mapping, and ETL development principles is needed. Experience with unit testing, data validation, and troubleshooting is required. Basic knowledge of layered data architectures (e.g., Bronze/Silver/Gold) is beneficial. Ability to work within established development frameworks and follow coding standards. Excellent analytical, problem-solving, and communication skills are highly valued. A Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent experience is required.
NTT DATA
IT Consulting