Infrastructure SRE - HPC
Sarvam AI
Sarvam AI
Sarvam is at the forefront of building India's sovereign AI platform, focusing on research, models, infrastructure, and applications to make AI impactful for the nation. We collaborate with leading enterprises and public institutions, backed by prominent investors and partnering with major Indian brands.
This role is crucial for managing our extensive, multi-vendor GPU fleet, which supports demanding training jobs and high-availability inference services. It requires specialized expertise in reliability engineering to ensure seamless operation across complex infrastructure challenges, going beyond standard Kubernetes administration.
We are assembling a team of specialists, valuing deep expertise in one area combined with practical knowledge across others. This collaborative approach ensures we can effectively manage and troubleshoot our shared fleet.
Operate the end-to-end GPU fleet, encompassing provisioning, observability, capacity management, and overall fleet health for both training and serving workloads.
Participate actively in an on-call rotation, authoring robust runbooks and conducting postmortems to implement lasting solutions.
Develop and maintain internal tooling essential for the team's operational success, rather than solely relying on off-the-shelf systems.
Collaborate closely with ML and platform teams to ensure the stability of large-scale runs and the predictability of serving latency.
A minimum of 5 years in infrastructure or site reliability engineering, with at least 2 years specifically managing large-scale GPU clusters.
Proven experience in on-call ownership of critical infrastructure, demonstrated through postmortems that resulted in tangible improvements.
Proficiency in Python or Go for building and maintaining internal tooling.
Working knowledge across five key areas of focus, enabling effective recognition, triage, and routing of issues outside your primary specialty.
Depth in one of the following areas: distributed high-performance storage (Lustre, GPFS, WEKA, or BeeGFS), fabric & RDMA networking (InfiniBand, RoCE, NVLink/NVSwitch), GPU systems reliability (NCCL, driver/firmware lifecycle, DCGM), Kubernetes platform reliability (GPU operator, scheduling, multi-tenancy), or training & inference workload reliability (hang detection, checkpoint/restart, HA/DR).
Bonus points for experience with Slurm and Kubernetes hybrid environments, on-premise GPU deployments, Indian NCPs, or multi-tenant GPU isolation.
Sarvam AI
IT Consulting