Data Engineer Sr (Databricks)
NTT DATA
NTT DATA
Seeking a Senior Data Engineer with extensive experience in Databricks to lead the replication and enhancement of data pipelines. This role involves migrating complex legacy data models into modern cloud-based architectures.
You will leverage Databricks and Delta Lake to build robust, scalable data solutions. The position offers a challenging opportunity to work with advanced data engineering concepts and integrate with a broad ecosystem of cloud services.
Join a dynamic team focused on transforming data infrastructure and ensuring data integrity throughout the migration process.
Develop and optimize Databricks notebooks, jobs, and workflows for data transformation. Implement Delta Lake tables with a Bronze/Silver/Gold architecture, ensuring ACID compliance and efficient data handling. Integrate Databricks with cloud platforms like AWS S3 or Azure ADLS, along with services such as Synapse and Key Vault. Optimize cluster performance and query efficiency to manage costs effectively. Build incremental load processes, CDC patterns, and batch schedules for large datasets. Collaborate with Snowflake and dbt teams to maintain consistent data models and contracts. Perform data validation and reconciliation between legacy systems and Databricks outputs. Adhere to coding standards, utilize version control (Git/Azure DevOps), and implement CI/CD practices. Provide support for defect fixes and post-implementation stabilization.
Requires 8+ years of experience with Databricks and a deep understanding of complex legacy data models (DB2/AS400, Guidewire). Proficiency in extracting data from DB2/AS400 using CDC or batch processes, working with JDBC connections and re-platforming SQL. Experience integrating with Guidewire Cloud Data Access (CDA) or InsuranceSuite to replicate P&C insurance schemas. Proven ability to transition legacy data stores to a scalable Medallion Architecture (Bronze, Silver, Gold). Expertise in Delta Lake optimization, building ETL/ELT pipelines with Apache Spark, and managing schema evolution and SCD Type 2. Skilled in refactoring legacy procedural code into scalable distributed patterns using PySpark, Spark SQL, and Scala. Experience implementing data governance, lineage tracing, and table-level security with Unity Catalog. Ability to automate data validation frameworks to ensure data quality during system transitions. Familiarity with Databricks serverless compute, optimal cluster sizing, and cost reduction strategies. Experience building event-driven pipelines with Databricks Auto Loader for both batch and streaming data ingestion.
NTT DATA
IT Consulting