ML Engineer (Training Infra), Foundational Models
Sarvam AI
Sarvam AI
Join Sarvam as an ML Engineer focused on Foundational Models Training Infrastructure and help build India's sovereign AI platform. This role is central to the development of our next generation of foundational models, demanding expertise in large-scale distributed systems and deep learning training.
You'll be instrumental in optimizing the performance, throughput, and stability of our training infrastructure. This is a critical systems role where your contributions directly impact the success and cost-efficiency of frontier training runs.
Own and advance the distributed training stack across extensive GPU clusters. Design and implement advanced parallelism strategies (data, tensor, pipeline, etc.) tailored to specific architectures and scales. Optimize end-to-end training throughput by fine-tuning kernel performance, communication, memory usage, and data loading. Develop and tune custom GPU kernels (CUDA, Triton) to maximize performance. Ensure the reliability of long-running training jobs through fault tolerance, deterministic restarts, and automated issue detection.
Possess a BS or MS in Computer Science or a related technical field, or equivalent practical experience. Require at least 3 years of experience in building ML training infrastructure or large-scale distributed systems. Demonstrate hands-on experience training large models using distributed frameworks like Megatron-LM, DeepSpeed, FSDP, or NeMo, including on-call experience for pretraining. Exhibit deep knowledge of GPU architecture, CUDA, and GPU profiling tools (Nsight, PyTorch profiler). Strong understanding of PyTorch internals is essential, with comfort in modifying low-level training code.
Sarvam AI
IT Consulting