Data Platform Modernization Engineer - Junior (Databricks)
NTT DATA
NTT DATA
NTT DATA is seeking a motivated Junior Data Platform Modernization Engineer with expertise in Databricks to join our dynamic team in Bangalore. This role is ideal for individuals passionate about innovation and growth within a forward-thinking organization.
As a key member of our team, you will contribute to transforming data platforms by migrating existing ETL workloads to the Databricks Lakehouse. You'll work with cutting-edge technologies to build robust data pipelines and implement efficient data transformation strategies.
Analyze and understand existing Informatica workflows and AWS Glue jobs, focusing on source-to-target mappings and transformation logic. Develop advanced Databricks notebooks and data pipelines to migrate current ETL processes. Implement comprehensive data transformations utilizing Python, PySpark, Spark SQL, and Databricks SQL. Construct Bronze, Silver, and Gold data layers according to established project architecture and coding standards. Design and build both batch and incremental processing pipelines leveraging Delta Lake and sophisticated watermark/control mechanisms. Develop and maintain Databricks Jobs/Workflows and Lakeflow Declarative Pipelines as needed. Integrate and reuse common project frameworks for logging, auditing, data quality, error handling, and notifications. Accurately convert Informatica transformations and business rules into equivalent Databricks implementations, ensuring preservation of business logic. Support the migration of AWS Glue jobs and database-centric processing into the Databricks environment. Conduct thorough unit testing and perform source-to-target data reconciliation for all migrated pipelines. Troubleshoot and resolve complex issues related to data, SQL, PySpark, and pipeline execution. Assist with integration testing, regression testing, UAT, and production validation phases. Collaborate closely with senior developers and technical leads to address technical challenges and dependencies. Adhere strictly to established coding, performance, security, and development standards. Actively participate in code reviews and integrate feedback to enhance code quality. Prepare detailed technical documentation and facilitate knowledge transfer for migrated workloads. Provide essential support for deployment, monitoring, production stabilization, and cutover activities.
A minimum of 3 to 6 years of experience in data engineering, ETL development, or application/data platform development is required. Demonstrated hands-on experience with Databricks and Apache Spark is essential. Strong proficiency in development using Python/PySpark and SQL is a must. Familiarity with Databricks notebooks, Spark SQL/Databricks SQL, Delta Lake, and Databricks Jobs/Workflows is necessary. Proven experience in developing batch or incremental ETL/ELT pipelines. Experience with at least one ETL technology such as Informatica PowerCenter, AWS Glue, SSIS, or similar is expected. Proficiency in working with relational databases and performing SQL-based data processing. Solid understanding of data transformation, source-to-target mapping, and ETL development principles. Experience with unit testing, data validation techniques, and effective troubleshooting methodologies. Basic knowledge of layered data architectures like Bronze/Silver/Gold is beneficial. Ability to work effectively within an established development framework and adhere to coding standards. Excellent analytical, problem-solving, and communication skills are highly valued. A Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent practical experience is required.
NTT DATA
IT Consulting