Performance Engineer, Inference
Sarvam AI
Sarvam AI
Join Sarvam's dedicated Performance Engineering team, focusing on optimizing inference for advanced AI models. This role is critical for ensuring efficient and cost-effective serving of diverse model families, including LLMs, ASR, TTS, and multimodal models, across a large-scale fleet of GPUs. You will be instrumental in defining the performance benchmarks that guide company-wide strategies.
This position is ideal for an experienced engineer passionate about pushing the boundaries of AI inference performance. You will operate at the forefront of serving technology, collaborating closely with SRE, model developers, and kernel engineers to achieve ambitious performance targets.
Own the end-to-end production serving path for large, distributed AI models. Integrate and optimize model and kernel artifacts into a robust multi-node, multi-tenant serving stack. Develop and train custom speculators, going beyond stock implementations to enhance performance. Proactively identify and address performance bottlenecks in areas such as KV cache, parallelism strategies, and routing.
Defend and improve key performance metrics including TTFT (p50/p95/p99), TPOT, throughput, GPU utilization, and cost per token. Conduct deep-dive profiling using tools like Nsight Systems and framework tracing to diagnose and resolve complex issues. Participate in architectural co-design and on-call rotations to maintain inference SLOs.
Requires a minimum of 5 years in ML systems, with at least 2 years specifically in production-scale inference serving. Demonstrated success in serving 100B+ parameter models across multi-node parallelism strategies is essential.
Proficiency in at least one of SGLang, vLLM, NVIDIA Dynamo, or TensorRT-LLM at a source-code modification level is required, along with familiarity with the others. Deep expertise in distributed serving concepts like disaggregated prefill-decode and distributed KV/cache transfer is a must.
Experience building and training speculative decoding models, coupled with a strong understanding of KV cache internals and various parallelism techniques (TP/PP/EP), is crucial. Fluency in C++ and CUDA for code modification, alongside profiling tools like Nsight Systems, is expected.
Sarvam AI
IT Consulting