Performance Engineer, Kernels
Sarvam AI
Sarvam AI
Join Sarvam's Performance Engineering team as a specialized Kernels Performance Engineer. This role is integral to optimizing our GPU code at the microsecond level, ensuring top-tier performance for a diverse range of model families including LLMs, Mixture-of-Experts, and multimodal systems.
You'll operate at the cutting edge of machine learning infrastructure, working with multi-node, multi-tenant fleets of H100, H200, and B200 GPUs. The Performance Engineering team is crucial for setting performance benchmarks and cost efficiency targets that drive the entire company's roadmap. This position offers a unique opportunity to impact system-level performance, cost, and GPU utilization.
If you possess deep expertise in kernel development and have a proven track record of shipping performance-critical code, we encourage you to apply. This is a high-impact role where your contributions directly translate to measurable improvements in production environments.
Take ownership of the kernel layer, developing custom CUDA, DSL-based, and PTX kernels to surpass the performance of standard libraries like cuBLAS, cuDNN, and FlashAttention.
Ship kernels that demonstrably outperform published baselines on real-world workloads. Your work will directly influence production p99 latency metrics.
Analyze and optimize kernel performance using tools such as Nsight Compute and Systems, identifying bottlenecks and implementing effective solutions.
Deeply understand and modify frameworks like CUTLASS and CuTe DSL, including their layout algebra.
Debug and modify PTX code to achieve compiler-level optimizations where necessary.
Gain expertise in multi-architecture differences, adapting kernel performance across Hopper, Blackwell, and future architectures.
A minimum of 5 years of experience in ML systems, with at least 2 years focused on authoring production-level CUDA kernels.
Demonstrated success in shipping a kernel that provided a measurable performance improvement over existing solutions.
Proficiency in CUDA kernel development, including advanced techniques like thread-block sizing, shared-memory layout, warp primitives, asynchronous copies (cp.async, TMA), and MMA selection.
Experience with CUTLASS and CuTe DSL at a modify-and-extend level, including a solid grasp of layout algebra.
Ability to debug and modify PTX code for compiler-level optimizations.
Fluency with Nsight Compute and Nsight Systems for performance analysis and tuning.
Experience authoring or modifying attention kernels such as FlashAttention-family, paged attention, MLA, sliding-window, or sparse attention.
Understanding of multi-architecture considerations, particularly how performance characteristics change between GPU generations like Hopper and Blackwell (e.g., TMA, WGMMA, tcgen05).
Sarvam AI
IT Consulting