Fabric Data Engineer - Pyspark

NTT DATA

Fresher Bengaluru Full Time Remote
NTT DATA logo
Posted : today
Actively hiring

Job description

Join NTT DATA as a Fabric Data Engineer specializing in PySpark. This role is integral to our growing Data & Analytics organization. We seek passionate individuals eager to innovate and grow within a forward-thinking company. You will be instrumental in designing, developing, and optimizing scalable data solutions.

Our team thrives on collaboration, working alongside architects, developers, and stakeholders to build robust enterprise data platforms. We are committed to fostering an inclusive and adaptable work environment where your contributions are valued and drive client success. If you are ready to shape the future of data engineering, we encourage you to apply.

Responsibilities

Key responsibilities include designing, developing, and maintaining enterprise data solutions using Microsoft Azure, with a focus on Azure Synapse Analytics and Microsoft Fabric. You will build scalable data ingestion, transformation, and processing solutions, optimizing data pipelines using Azure Data Factory and Fabric Data Factory.

This role involves designing solutions that integrate data from diverse sources, developing reusable data engineering frameworks, and implementing both batch and near-real-time processing. Monitoring, troubleshooting, and optimizing Azure data workloads are crucial. You will also support the modernization and migration of legacy data platforms to Azure.

Responsibilities extend to designing and developing data warehouses and Lakehouse solutions, working with Synapse Pipelines and Fabric Data Factory. Developing notebooks using Python, PySpark, and Spark SQL, along with building and optimizing data transformation processes, is essential. You will collaborate on data integration, data lake, lakehouse, and data warehouse architectures, ensuring secure and scalable cloud data solutions.

Qualifications

Ideal candidates possess significant hands-on experience with Microsoft Azure data services, Azure Synapse Analytics, and/or Microsoft Fabric. A strong background in data engineering, data integration, data transformation, and enterprise data platforms is required.

Proficiency in PySpark, Python, and Spark SQL for developing scalable data transformations and processing large datasets is essential. Experience with ETL/ELT pipeline design, ingesting various data types from multiple enterprise sources, and implementing incremental data loads and CDC strategies is expected.

Familiarity with Azure Data Lake Storage Gen2, Microsoft OneLake, and implementing data layers (Bronze, Silver, Gold) is necessary. Experience with dimensional data modeling (Star/Snowflake schemas) and implementing Slowly Changing Dimensions (SCD) is also important. You should be adept at working with Delta Lake and related technologies, along with partitioning and performance optimization strategies.

Essential Skills

PySparkSQLMicrosoft AzureAzure Synapse AnalyticsMicrosoft FabricData EngineeringData IntegrationData TransformationPythonSpark SQLETLELTData WarehousingData LakeLakehouseAzure Data FactoryAzure Data Lake Storage Gen2Azure SQL DatabaseAzure FunctionsAzure Event HubsMicrosoft Entra IDDelta Lake

Good to Have

Power BIDevOps

Highlights

  • Actively hiring

More Details

RoleFabric Data Engineer - Pyspark
IndustryInformation Technology & Services
DepartmentData Science
Employment TypeFull Time, Remote

About the Company

NTT DATA logo

NTT DATA

Information Technology & Services

Fabric Data Engineer - Pyspark at NTT DATA | SkillMX | SkillMX