Data Engineer - Python AND Kafka AND (Hadoop OR HDFS OR Hive) AND Snowflake AND apache AND (iceberg

NTT DATA

3–5 yrs Bengaluru Full Time Hybrid (office + remote)
NTT DATA logo
Posted : today
Actively hiring

Job description

Join NTT DATA and contribute to a vital datastore migration project, transforming an on-premise Data Lake into an AWS-hosted LakeHouse. This role is integral to ensuring data integrity and functionality during a high-visibility migration for Goldman Sachs.

This position involves refactoring and migrating extraction logic, job scheduling, and physical datasets, while meticulously validating data accuracy. You will also be responsible for translating and optimizing legacy consumption patterns to ensure seamless compatibility with Snowflake and Iceberg, driving the delivery of essential data products.

Responsibilities

- Refactor and migrate extraction logic and job scheduling from legacy systems to the new Lakehouse environment. - Execute the physical migration of datasets, ensuring complete data integrity throughout the process. - Serve as a technical liaison, engaging with internal clients and data owners to facilitate handoffs and obtain sign-offs for migrated assets. - Translate and optimize legacy SQL and Spark-based consumption patterns for Snowflake and Iceberg compatibility. - Analyze usage patterns to deliver required data products effectively. - Implement rigorous data validation using reconciliation frameworks to confirm functional equivalence of migrated data with production flows.

Qualifications

- Bachelor's or Master's degree in Computer Science, Applied Mathematics, Engineering, or a related quantitative field. - Minimum of 3-5 years of professional, hands-on coding experience in a collaborative team environment. - Professional proficiency in Python or Java, with basic scripting and SQL troubleshooting abilities. - Deep familiarity with the Software Development Life Cycle (SDLC) and CI/CD best practices, including K8s deployment experience. - Sophisticated understanding of temporal data modeling, schema evolution (Apache Iceberg), performance optimization (partitioning and clustering), and architectural theory (normalization vs. denormalization, keys). - Experience with Kafka, ANSI SQL, FTP, Apache Spark, JSON, Avro, Parquet, Hadoop (HDFS/Hive), Snowflake, and Apache Iceberg is required for the collective team. Aptitude for learning new workflows and language constructs is essential.

Essential Skills

PythonKafkaHadoopHDFSHiveSnowflakeApacheIcebergSQLSparkJavaSDLCCI/CDK8sTemporal Data ModelingSchema ManagementPerformance OptimizationJSONAvroParquet

Highlights

  • Actively hiring

More Details

RoleData Engineer - Python AND Kafka AND (Hadoop OR HDFS OR Hive) AND Snowflake AND apache AND (iceberg
Employment TypeFull Time, Hybrid (office + remote)

About the Company

NTT DATA logo

NTT DATA

IT Consulting

Data Engineer - Python AND Kafka AND (Hadoop OR HDFS OR Hive) AND Snowflake AND apache AND (iceberg at NTT DATA | SkillMX | SkillMX