ETL Data - Senior Engineer (Noida, UP, India)
Iris Software
Iris Software
Join a leading IT services company recognized as one of India's Top 25 Best Workplaces in the IT industry. At Iris Software, we empower our employees to achieve their career best while working in an award-winning, growth-oriented culture. We are a fast-growing organization committed to being our clients' most trusted technology partner and a premier destination for top industry talent.
Our vision is to enable enterprise clients across financial services, healthcare, transportation & logistics, and professional services to thrive through technology-enabled transformation. We specialize in complex, mission-critical applications leveraging cutting-edge technologies in application and product engineering, data & analytics, cloud, DevOps, MLOps, quality engineering, and business automation. We foster an environment where ownership, growth, and impact are paramount, supported by personalized career development, continuous learning, and mentorship.
This role involves designing and implementing scalable data engineering solutions using PySpark and modern distributed data processing frameworks. Key responsibilities include defining data ingestion, transformation, and processing architectures aligned with business objectives, and optimizing Snowflake or Delta Lake on Databricks solutions for enterprise-scale data platforms.
You will lead the implementation of high-performance batch and streaming data pipelines, defining standards for data streaming, integration frameworks, and scalable processing patterns. The role requires architecting workflow orchestration solutions using Apache Airflow or Databricks Workflows and establishing robust monitoring, scheduling, and operational controls for reliable pipeline execution.
Emphasis will be placed on driving data quality, validation, reconciliation, and governance practices. You will design solutions based on modern Lakehouse architecture principles, data observability, and platform engineering standards to enhance scalability and operational visibility. Promoting responsible AI-assisted engineering and ensuring adherence to engineering, scalability, and performance standards through rigorous review of pipeline designs are also key aspects of this position.
We are seeking a Senior Engineer with mandatory skills in PySpark, Databricks Workflows, Delta Lake on Databricks, and Amazon Kinesis. A strong understanding of Big Data technologies, particularly Apache Spark and Python, is essential.
Proficiency in Data Quality & Validation, and database programming including SQL, is required. Experience with cloud platforms, specifically AWS services such as SNS, SQS, Kinesis, CloudWatch, S3, IAM, Secrets Manager, KMS, Cognito, API Gateway, Glue, EMR, Redshift, Dynamo DB, Aurora, RDS, and EC2, is crucial. Familiarity with ETL & Data Integration processes and middleware like API (SOAP, REST) is also expected.
Candidates should possess strong behavioral competencies including ownership, collaboration, quality-focused engineering, analytical thinking, adaptability, effective communication, and meticulous attention to detail. A commitment to continuous improvement, knowledge sharing, mentoring, and balancing scalability, performance, reliability, and business priorities is highly valued. Experience with Unix/Linux Shell scripting is also a requirement.
Iris Software
Information Technology & Services