Researcher, Vision
Sarvam AI
Sarvam AI
Join Sarvam, a pioneering company dedicated to building India's sovereign AI platform. We are at the forefront of AI innovation, developing end-to-end solutions that cater specifically to India's needs. Our work spans research, model development, infrastructure, and applications, impacting diverse sectors through collaborations with leading enterprises and public institutions. Backed by prominent investors and partnering with major Indian brands, Sarvam offers a unique opportunity to shape the future of AI in India.
As a Researcher in Vision, you will be instrumental in the entire lifecycle of vision-language model (VLM) development, from data curation and training to evaluation and deployment. Your responsibilities will include:
- Researching and advancing vision-language architectures, including encoders, fusion mechanisms, and pretraining objectives. - Developing innovative training methodologies such as SFT and RLHF, tailored for multilingual VLMs. - Investigating and optimizing data strategies, focusing on mixture compositions, quality signals, and synthetic data generation. - Building robust evaluation frameworks and benchmarks, with a special emphasis on Indic multimodal tasks. - Analyzing model failure modes, enhancing robustness, and exploring interpretability. - Collaborating closely with engineering teams to ensure research is scalable and testable. - Contributing to the broader research community through open-source initiatives and collaborations.
We are seeking individuals with a profound understanding of vision-language models, encompassing training dynamics, architectural trade-offs, and failure modes. A strong track record of impactful research, demonstrated through publications, technical reports, or successful project implementations, is essential.
Key qualifications include:
- Rigorous experimental design capabilities for isolating variables and deriving defensible conclusions. - Proficiency in PyTorch for end-to-end experiment execution. - Intellectual breadth to tackle diverse problems across data, training, and evaluation.
Preferred qualifications:
- A PhD or Master's degree with relevant research experience in Machine Learning, Computer Vision, NLP, or a related domain. - Publications at top-tier (A/A*) academic venues. - Experience with multilingual or low-resource language modeling. - Familiarity with document understanding, OCR, or structured visual prediction tasks. - Experience in large-scale data curation and its impact on model quality.
Sarvam AI
AI / Machine Learning