Performance Engineer, On-Device Inference
Sarvam AI
Sarvam AI
Sarvam is at the forefront of developing India's sovereign AI platform, encompassing research, models, infrastructure, and applications. Our mission is to create AI solutions that genuinely benefit India. We collaborate with leading enterprises and public institutions, supported by prominent investors and partnering with major Indian brands.
This role involves optimizing Sarvam's AI models for production readiness across diverse chipsets like Intel xPU, ARM xPU, Apple xPU, and Nvidia/AMD GPUs. You will manage 1-2 model and chipset pairs from development through integration, ensuring end-to-end quality. Key tasks include quantization, accuracy validation, performance benchmarking, and documentation, along with authoring deployment guides.
Candidates should possess at least 3 years of experience in ML systems and have strong PyTorch and ONNX export skills, including navigating complexities like dynamic shapes and control flow. Proven experience with production-level quantization on real models is essential. Proficiency with at least two of the following runtimes is required: ONNX Runtime, TensorRT, CoreML, OpenVINO, QNN, or LiteRT. Demonstrating profiling fluency on at least one platform is also a must.
Sarvam AI
AI / Machine Learning