Data Scientist - Evaluations, Chanakya
Sarvam AI
Sarvam AI
Join Sarvam, a pioneering company building India's sovereign AI platform. This role focuses on anchoring the evaluations function, developing innovative frameworks to assess model and system quality in real-world, high-stakes applications. You will be instrumental in ensuring our AI solutions are robust and reliable for critical use cases.
This position offers the opportunity to work closely with MLOps engineers, product managers, and deployment teams. Your evaluations will directly influence deployment decisions and ongoing performance monitoring, making your contribution vital to our mission of making AI work for India.
Design and implement bespoke evaluation frameworks for diverse AI applications including document comprehension, command summarization, and geospatial reasoning.
Collaborate with domain experts to define quality metrics and translate operational needs into measurable signals.
Conduct rigorous evaluation cycles both pre- and post-deployment, creating dashboards to visualize model quality in production environments.
Proactively identify potential failure modes, edge cases, and distribution shifts in AI outputs.
Partner with the MLOps Engineer to automate and operationalize evaluation pipelines, ensuring they are versioned and reproducible.
Develop and manage domain-specific datasets for fine-tuning, evaluation, and benchmarking, including overseeing human annotation workflows.
Communicate findings and quality reports to inform the product and engineering roadmap.
Possess 3–6 years of experience in data science, ML research, or applied AI, with a minimum of 2 years specifically working with LLMs in production. A strong foundation in statistics and probability is essential for designing valid evaluations.
Demonstrated experience in building evaluation frameworks from the ground up, including defining custom metrics, assessing inter-rater reliability, and employing red-teaming methodologies.
Proficiency in Python, coupled with practical experience using libraries like pandas and NumPy, and familiarity with tools such as HuggingFace datasets, RAGAS, EleutherAI Eval Harness, or LangSmith is required.
Experience with prompt engineering, model fine-tuning, or Reinforcement Learning from Human Feedback (RLHF) in applied settings is beneficial.
Ability to effectively work with unstructured domain data, including PDFs, doctrine documents, transcripts, and field reports. Experience in high-stakes domains like healthcare, legal, defense, or finance is a plus.
Sarvam AI
AI / Machine Learning