Platform Engineer - AI Infrastructure
Sarvam AI
Sarvam AI
Join Sarvam in building India's sovereign AI platform, driving innovation across models, infrastructure, and applications. We focus on making AI work for India, partnering with leading enterprises and public institutions. This role is crucial in developing the sophisticated platform that manages our extensive GPU fleet for both demanding training and high-performance inference workloads.
As a Platform Engineer, you'll be instrumental in designing and implementing the core systems that enable seamless GPU utilization by our ML teams. Your work will focus on the build side, creating robust scheduling, scaling, multi-tenancy, and serving layers. Treat our platform as a product, with internal ML engineers as your valued customers.
Develop and ship control-plane services, scheduler integrations, and autoscaling controllers. Build the inference-serving platform, RBAC, and quota systems. Create observability and cost tooling, along with the CLI and APIs for ML engineers. Design and implement the serving platform for scalable, multi-tenant endpoints, including intelligent routing and rollout machinery. Develop the elasticity layer for both training and serving workloads, managing capacity pooling and burst handling. Enhance scheduling and orchestration capabilities, focusing on gang scheduling, priority, preemption, and topology-aware placement. Implement robust multi-tenancy, RBAC, and isolation mechanisms, including secrets management and audit logging. Develop networking components, CNI configuration, and multi-cluster connectivity solutions. Build the storage and data path abstractions for efficient data placement. Craft the developer experience through CLIs, SDKs, and APIs for self-service job submission. Develop the infrastructure-as-code control plane for reproducible cluster provisioning.
A minimum of 5 years of experience building production-grade infrastructure or platform software, with a proven track record of shipping services and control planes. Strong software engineering skills in Go or Python, demonstrating the ability to build and debug maintainable systems. Deep understanding of Kubernetes internals, including controller development, scheduler operations, and API machinery. Working knowledge of GPU-specific platform constraints such as MIG and GPU sharing, gang scheduling, and topology-aware placement. A product-oriented mindset focused on delivering user-friendly APIs and abstractions for internal teams. Capability to own a feature end-to-end, from initial design through deployment and comprehensive documentation.
Sarvam AI
AI / Machine Learning