Embedded Data Scientist, Chanakya
Sarvam AI
Sarvam AI
Join Sarvam, a pioneering company building Sovereign AI for India. We are creating a full-stack AI platform, encompassing research, models, infrastructure, and applications, with the core mission of making AI truly impactful for India. Partnering with leading enterprises and public institutions, we are backed by prominent venture capital firms and collaborate with India's top brands.
This role involves transforming complex client data into formats that AI systems can effectively process. You'll be embedded alongside Strategic Deployment Engineers at client sites, directly engaging with their data environments to understand, structure, and operationalize large-scale datasets. This includes working with diverse data types such as documents, images, audio, geospatial data, and structured records, and designing the semantic structures essential for AI interpretation and reasoning.
Your responsibilities will include defining how data is represented within AI systems, such as document segmentation, metadata definition, and entity/relationship representation. You will design ontologies, tagging systems, and knowledge graph structures to optimize the reasoning engine's performance. Often, you'll work with sensitive datasets in environments that may lack standard tooling. You will be accountable for the data layer's quality, ensuring it supports reliable, large-scale reasoning.
Key duties involve understanding client data landscapes, designing domain ontologies, and defining document segmentation strategies. You will work with heterogeneous datasets, determining indexing and linking methods, and collaborating with engineers to build data ingestion pipelines. Evaluating AI system performance and refining structures for improvement will also be crucial, alongside defining benchmarks and translating data insights into actionable signals for product and engineering teams.
We are seeking candidates with 2-5 years of experience in data science, applied machine learning, or large-scale data analysis. Proficiency in Python, including pandas and NumPy, along with modern NLP or LLM tools, is essential. A solid understanding of ML fundamentals is required to contribute to evaluation design and collaborate with the models team. Experience with large unstructured datasets like documents and reports, as well as familiarity with LLM-based systems and retrieval pipelines, is highly valued.
We look for individuals who have successfully built rigorous systems from messy, unstructured real-world data and are adept at designing structure where it doesn't exist. The ability to translate complex data insights into clear explanations for engineers and stakeholders is key. You should be comfortable working autonomously in client environments, moving fluidly between domain understanding, data modeling, and AI system design, and bridging the technical with the operational context.
Sarvam AI
AI / Machine Learning