Data Platform Modernization Engineer - Junior (Databricks)

NTT DATA

3–6 yrs Bengaluru Full Time Hybrid (office + remote)
NTT DATA logo
Posted : 1 week ago
Actively hiring

Job description

Join NTT DATA, a global leader in business and technology services, and contribute to our mission of accelerating client success through responsible innovation. We are seeking a Junior Data Platform Modernization Engineer specializing in Databricks to enhance our team in Bangalore, India. This role offers a fantastic opportunity to work with cutting-edge technologies and drive significant impact within a forward-thinking organization.

This position is ideal for individuals passionate about data engineering and eager to grow their skills in a dynamic environment. You will play a key role in transforming existing data pipelines and modernizing our data platform, contributing to our enterprise-scale AI and digital infrastructure capabilities.

Responsibilities

Analyze existing Informatica workflows and AWS Glue jobs to grasp source-to-target mappings and transformation logic. Develop Databricks notebooks and data pipelines for migrating ETL workloads to the Databricks Lakehouse. Craft data transformations utilizing Python, PySpark, Spark SQL, and Databricks SQL. Implement Bronze, Silver, and Gold layer processing following established architectural patterns and coding standards. Build robust batch and incremental processing pipelines using Delta Lake and advanced watermark/control mechanisms. Develop and maintain Databricks Jobs/Workflows and Lakeflow Declarative Pipelines as needed. Leverage common project frameworks for logging, auditing, data quality, error handling, and notifications. Convert Informatica transformations, mappings, lookups, and business rules into equivalent Databricks implementations, ensuring business logic preservation. Support the migration of AWS Glue jobs and database-based processing into Databricks. Conduct unit testing and perform source-to-target data reconciliation for all migrated pipelines. Investigate and resolve issues related to data, SQL, PySpark, and pipeline execution. Assist with integration testing, regression testing, UAT, and production validation activities. Collaborate with senior developers and technical leads to resolve technical challenges and dependencies. Adhere to established coding, performance, security, and development standards. Participate actively in code reviews and integrate feedback effectively. Prepare comprehensive technical documentation and facilitate knowledge transfer for migrated workloads. Provide support for deployment, monitoring, production stabilization, and cutover procedures.

Qualifications

Possess 3–6 years of experience in data engineering, ETL development, or application/data platform development. Demonstrate hands-on expertise with Databricks and Apache Spark. Exhibit strong development proficiency in Python/PySpark and SQL. Showcase experience with: - Databricks notebooks - Spark SQL / Databricks SQL - Delta Lake - Databricks Jobs/Workflows Proven experience developing batch or incremental ETL/ELT pipelines. Familiarity with at least one ETL technology like Informatica PowerCenter, AWS Glue, SSIS, or similar. Experience working with relational databases and SQL-based data processing. Solid understanding of data transformation, source-to-target mapping, and ETL development principles. Experience with unit testing, data validation, and troubleshooting techniques. Basic understanding of layered data architectures such as Bronze/Silver/Gold. Ability to operate effectively within established development frameworks and adhere to coding standards. Possess good analytical, problem-solving, and communication skills. Hold a Bachelor's degree in Computer Science, Information Technology, Engineering, or possess equivalent relevant experience.

Essential Skills

DatabricksApache SparkPythonPySparkSQLSpark SQLDatabricks SQLDelta LakeETLELTData TransformationSource-to-Target MappingUnit TestingData ValidationTroubleshootingRelational Databases

Good to Have

Databricks CertificationLakeflow Declarative PipelinesDelta Live TablesUnity CatalogDatabricks Asset BundlesCI/CDAWS S3AWS GlueAWS RedshiftAuto LoaderStreaming IngestionCDCSCD Type 1SCD Type 2Data Quality FrameworksAudit FrameworksError Handling FrameworksInformatica PowerCenterAWS LambdaSpark Performance OptimizationDelta Lake Performance OptimizationCloud Data ModernizationData Migration

Highlights

  • Actively hiring

More Details

RoleData Platform Modernization Engineer - Junior (Databricks)
IndustryInformation Technology & Services
DepartmentData Engineering, Software Development
Employment TypeFull Time, Hybrid (office + remote)

About the Company

NTT DATA logo

NTT DATA

Information Technology & Services

Data Platform Modernization Engineer - Junior (Databricks) at NTT DATA | SkillMX | SkillMX