AWS Data Engineer (Noida, UP, India)
Iris Software
Iris Software
Join Iris Software, recognized as one of India's Top 25 Best Workplaces in the IT industry, and contribute to your best work. As a rapidly growing IT services company, we empower our associates to shape their own success stories.
Our vision is to be a premier technology partner, providing a platform for top industry professionals to reach their full potential. With a global presence across India, the U.S.A., and Canada, we support enterprise clients with technology-driven transformations in financial services, healthcare, transportation & logistics, and professional services.
We specialize in complex, mission-critical applications leveraging cutting-edge technologies such as Application & Product Engineering, Data & Analytics, Cloud, DevOps, Data & MLOps, Quality Engineering, and Business Automation.
Define and implement enterprise data engineering strategies aligned with organizational goals and modernization efforts. Establish comprehensive data engineering standards, governance frameworks, and best practices. Lead the design of large-scale data processing architectures using PySpark and advanced data platform technologies. Develop enterprise standards for Snowflake and Delta Lake data platforms for analytical and operational workloads. Drive real-time and event-driven data architecture initiatives using technologies like Apache Kafka or Amazon Kinesis. Establish governance standards for data ingestion, transformation, streaming, and processing. Define workflow orchestration, scheduling, and operational governance using Apache Airflow or Databricks Workflows. Implement data quality, validation, monitoring, and operational excellence frameworks. Define enterprise standards for data products, metadata management, discoverability, and trusted business data consumption. Establish architecture standards for modern Lakehouse platforms, data observability, platform engineering, and scalable cloud-native data ecosystems. Define AI-ready data foundation strategies for structured and unstructured data, vector architectures, and future GenAI initiatives. Collaborate with business stakeholders to translate objectives into scalable data platform capabilities and architecture decisions. Promote AI-assisted engineering practices to enhance developer productivity and code quality. Lead architecture reviews to ensure solutions meet scalability, reliability, and performance requirements. Guide teams on distributed data processing, streaming architectures, and modern data platform best practices. Identify and mitigate platform risks, scalability bottlenecks, and architectural challenges. Align data initiatives with organizational objectives through collaboration with teams and leadership. Drive continuous improvement in platform maturity, engineering excellence, and delivery effectiveness.
Proven experience of 6-8 years in Data Engineering. Expertise in Apache Spark, PySpark, and SQL. Proficiency with Databricks Workflows and Delta Lake on Databricks. Experience with Amazon Kinesis for real-time data streaming. Familiarity with AWS services such as AWS Glue, AWS S3, AWS EBS, Amazon RDS, and DynamoDB. Experience with distributed data processing and streaming architectures. Strong understanding of data governance, quality, and operational excellence frameworks. Ability to define and implement enterprise-level data strategies and standards. Experience partnering with business stakeholders to translate requirements into technical solutions. Excellent leadership, communication, and strategic thinking skills. Demonstrated ability to drive initiatives, promote innovation, and ensure solution excellence.
Iris Software
IT Consulting