Fabric Data Engineer - Pyspark
NTT DATA
NTT DATA
Join NTT DATA as a Fabric Data Engineer specializing in PySpark. This role is integral to our growing Data & Analytics organization. We seek passionate individuals eager to innovate and grow within a forward-thinking company. You will be instrumental in designing, developing, and optimizing scalable data solutions.
Our team thrives on collaboration, working alongside architects, developers, and stakeholders to build robust enterprise data platforms. We are committed to fostering an inclusive and adaptable work environment where your contributions are valued and drive client success. If you are ready to shape the future of data engineering, we encourage you to apply.
Key responsibilities include designing, developing, and maintaining enterprise data solutions using Microsoft Azure, with a focus on Azure Synapse Analytics and Microsoft Fabric. You will build scalable data ingestion, transformation, and processing solutions, optimizing data pipelines using Azure Data Factory and Fabric Data Factory.
This role involves designing solutions that integrate data from diverse sources, developing reusable data engineering frameworks, and implementing both batch and near-real-time processing. Monitoring, troubleshooting, and optimizing Azure data workloads are crucial. You will also support the modernization and migration of legacy data platforms to Azure.
Responsibilities extend to designing and developing data warehouses and Lakehouse solutions, working with Synapse Pipelines and Fabric Data Factory. Developing notebooks using Python, PySpark, and Spark SQL, along with building and optimizing data transformation processes, is essential. You will collaborate on data integration, data lake, lakehouse, and data warehouse architectures, ensuring secure and scalable cloud data solutions.
Ideal candidates possess significant hands-on experience with Microsoft Azure data services, Azure Synapse Analytics, and/or Microsoft Fabric. A strong background in data engineering, data integration, data transformation, and enterprise data platforms is required.
Proficiency in PySpark, Python, and Spark SQL for developing scalable data transformations and processing large datasets is essential. Experience with ETL/ELT pipeline design, ingesting various data types from multiple enterprise sources, and implementing incremental data loads and CDC strategies is expected.
Familiarity with Azure Data Lake Storage Gen2, Microsoft OneLake, and implementing data layers (Bronze, Silver, Gold) is necessary. Experience with dimensional data modeling (Star/Snowflake schemas) and implementing Slowly Changing Dimensions (SCD) is also important. You should be adept at working with Delta Lake and related technologies, along with partitioning and performance optimization strategies.
NTT DATA
Information Technology & Services