ML Engineer (Data), Foundational Models
Sarvam AI
Sarvam AI
Join Sarvam in building India's sovereign AI platform, focusing on foundational models. This role is pivotal in developing the data infrastructure that powers our next generation of AI models, ensuring petabyte-scale data curation, filtering, and quality assurance.
This position demands a strong engineering and research focus, delving into complex challenges such as large-scale deduplication, quality modeling, contamination detection, and mixture design. You will be instrumental in shaping the data that trains our advanced AI systems, treating data quality with the same scientific rigor as model architecture choices.
Design and implement robust, large-scale data pipelines for pre-training and post-training, covering ingestion, parsing, normalization, filtering, deduplication, tokenization, and packing at petabyte scale.
Develop and continuously enhance sophisticated quality filtering systems, incorporating model-based classifiers and advanced contamination detection techniques.
Spearhead data mixture design, curriculum development, and annealing strategies in close collaboration with the research team, ensuring precise traceability of data exposure for every model.
Build essential tooling to empower researchers and engineers in analyzing, slicing, attributing, and debugging datasets.
Scale the data pipeline to expertly handle diverse multilingual corpora, code, mathematical data, multi-source web content, and licensed datasets, while meticulously tracking provenance and licensing.
A Bachelor's or Master's degree in Computer Science or a related technical field, or equivalent practical experience.
A minimum of 3 years of experience in building extensive data systems, including petabyte-scale processing or distributed data pipelines. Exceptional candidates with a strong systems background and less experience may be considered.
Demonstrated hands-on expertise in data curation and filtering specifically for Large Language Model (LLM) training.
Profound understanding of distributed data processing frameworks such as Spark, Ray, Beam, or Dask, and their underlying storage systems.
Proficiency in Python, with a solid grasp of low-level data path components like tokenization, sharding, packing, and I/O patterns, including their performance implications.
Sarvam AI
AI / Machine Learning