Data Science Architect | 9 to 14 years | Gurgaon
Capgemini
Capgemini
Join Capgemini Engineering, a global leader in engineering services, and collaborate with a diverse team of engineers, scientists, and architects. We empower the world's most innovative companies to reach their full potential. From autonomous vehicles to advanced robotics, our experts in digital and software technologies push boundaries to deliver unique R&D and engineering services across all sectors. Discover a career brimming with opportunities for growth and impact, where each day presents new challenges and experiences.
This role involves owning the entire data science lifecycle for multi-industry projects. Responsibilities span from presales and opportunity definition to designing greenfield ML/AI platforms, developing and deploying models, and ensuring program governance throughout steady-state delivery. The vision is to move towards autonomous, self-optimizing AI/ML operations.
- Lead presales activities, including RFP/RFI responses, AI/ML solution architecture design, effort estimation, and commercial proposal development. Present strategic win themes to executive stakeholders. - Oversee the complete data science lifecycle: problem formulation, exploratory data analysis, feature engineering, model development, validation, and deployment. - Architect new data and ML platforms, encompassing cloud-native data lakes/lakehouses, feature stores, and MLOps pipelines. This includes selecting appropriate vendors and tools. - Define strategies for achieving autonomous, self-optimizing ML operations, incorporating automated retraining, drift detection, and closed-loop model monitoring. - Drive large-scale modernization of existing data and analytics platforms, managing migrations from legacy systems to target architectures with minimal business interruption. - Collaborate with data engineering teams to ensure pipeline quality, governance, and lineage, while personally managing model design, validation, and business impact assessment. - Provide delivery governance for multi-year, multi-workstream AI/ML transformation programs, overseeing scope, schedule, risk, quality, and financial aspects. - Serve as the primary technical point of contact across design, build, deploy, and operate phases, coordinating efforts among data engineering, MLOps, and business teams. - Establish and monitor model and program KPIs (e.g., accuracy, drift, business impact), and report findings to steering committees and executive dashboards. - Mentor data science and delivery leads, and develop reusable accelerators, ML frameworks, and best-practice playbooks for use across various engagements.
This role requires expertise in Data Science & ML, with a strong background in designing and deploying advanced analytics and AI solutions using traditional Machine Learning techniques. A deep understanding of statistical concepts is essential, covering areas such as probability theory, hypothesis testing, confidence intervals, Bayesian statistics, and experimental design.
Proficiency in data modeling is crucial, including the ability to evaluate data quality, identify bias and fairness concerns, perform causal inference, and develop explainable AI solutions. Experience with Big Data technologies like Spark, Hadoop, Databricks, and Kafka is necessary, alongside strong skills in Python and SQL.
Expertise in Deep Learning and Generative AI is also required, encompassing LLMs, RAG, prompt engineering, vector databases, and agentic AI frameworks. A demonstrated ability to translate complex business challenges into data-driven solutions is key, ensuring responsible AI practices, model governance, transparency, and measurable business outcomes.
Capgemini
Engineering