Senior Data Engineer
Iris Software
Iris Software
We are seeking a seasoned AI Data Engineer to architect and refine data pipelines that fuel AI/ML solutions. This role is integral to data collection through designing and implementing data instruments, managing aggregation and cleaning, and partnering with data scientists and ML engineers for interpretation and implementation. The position is vital for preparing high-quality, scalable, and compliant data for advanced analytics, automation, and machine learning use cases, while also supporting data visualization and trend analysis.
Iris Software is recognized as one of India's Top 25 Best Workplaces in the IT industry, offering an award-winning culture that values talent and ambition. As one of the fastest-growing IT services companies, we empower our associates to own and shape their career journeys. Our vision is to be our clients' most trusted technology partner and the premier choice for industry professionals seeking to realize their full potential.
With a global team across India, the U.S.A., and Canada, we support enterprise clients in transforming their operations across financial services, healthcare, transportation & logistics, and professional services. Our work encompasses complex, mission-critical applications leveraging cutting-edge technologies in Application & Product Engineering, Data & Analytics, Cloud, DevOps, Data & MLOps, Quality Engineering, and Business Automation.
Design and implement robust, scalable data pipelines to support AI/ML model development and deployment. Clean, transform, and curate structured and unstructured data from diverse sources to ensure model-ready quality. Collaborate with data scientists, ML engineers, and business teams to ensure data readiness, usability, and alignment with AI objectives. Develop and maintain metadata management, data lineage, and data quality frameworks to support AI governance and compliance. Enable advanced feature engineering capabilities and implement real-time data streaming solutions for AI applications. Design and deploy data collection instruments and oversee data aggregation processes. Ensure data compliance, privacy, and ethical use standards across all AI workflows. Support enterprise-wide data initiatives including business glossary development, taxonomy creation, and automation goals for the DARE program.
Proficiency in Python and advanced SQL is essential. Experience with Spark or other distributed processing frameworks is required. Must have experience with cloud platforms, specifically AWS. Familiarity with orchestration tools such as Airflow is necessary. Demonstrated ability in data pipeline automation is a key requirement. Solid understanding of data governance and data quality principles is mandatory. Experience with Kafka/Kinesis streaming, feature stores, LLM-ready data pipelines, and MLOps exposure are considered advantageous.
Iris Software
Information Technology & Services