Azure Subcontractor
Birlasoft
Birlasoft
Join our team as an experienced Databricks Developer focused on crafting and managing robust data pipelines and analytics solutions. You'll leverage your expertise in Apache Spark, Python, and SQL, alongside cloud data engineering technologies, to enhance enterprise data integration and analytics capabilities.
This role is crucial for building and maintaining scalable data infrastructure, ensuring efficient data processing and analysis for our key initiatives.
Key responsibilities include designing, developing, and maintaining ETL/ELT pipelines within Databricks. You will build scalable data processing frameworks using Apache Spark (PySpark) and develop both batch and streaming data pipelines. Integration of data from diverse sources, including databases, APIs, and cloud storage, is essential.
Furthermore, you will implement data transformation, cleansing, and validation processes, alongside developing and optimizing Delta Lake tables. Creating reusable Databricks notebooks, workflows, and jobs, while actively monitoring and troubleshooting pipeline performance, will be a core part of the role. Collaboration with stakeholders to understand data needs and adherence to coding standards and CI/CD best practices are also vital.
We are seeking a professional with strong experience in Databricks and a deep understanding of Apache Spark/PySpark. Proficiency in Python and strong SQL skills are mandatory. Experience with ETL/ELT development, Delta Lake, and data warehousing concepts is required. Familiarity with Git and version control best practices is also essential.
Preferred skills include experience with cloud platforms like Microsoft Azure (Azure Databricks, ADLS Gen2, Azure Data Factory), AWS (S3, Glue, EMR), or Google Cloud (BigQuery, Cloud Storage). Experience with streaming technologies such as Spark Structured Streaming and orchestration tools like Airflow is beneficial. A solid understanding of data governance and security concepts is also a plus.
Birlasoft
IT Consulting