Performance Engineer, On-Device Inference

Sarvam AI

3+ yrs Bengaluru Full Time Hybrid (office + remote)
Sarvam AI logo
Posted : yesterday
Actively hiring

Job description

Sarvam is at the forefront of developing India's sovereign AI platform, encompassing research, models, infrastructure, and applications. Our mission is to create AI solutions that genuinely benefit India. We collaborate with leading enterprises and public institutions, supported by prominent investors and partnering with major Indian brands.

Responsibilities

This role involves optimizing Sarvam's AI models for production readiness across diverse chipsets like Intel xPU, ARM xPU, Apple xPU, and Nvidia/AMD GPUs. You will manage 1-2 model and chipset pairs from development through integration, ensuring end-to-end quality. Key tasks include quantization, accuracy validation, performance benchmarking, and documentation, along with authoring deployment guides.

Qualifications

Candidates should possess at least 3 years of experience in ML systems and have strong PyTorch and ONNX export skills, including navigating complexities like dynamic shapes and control flow. Proven experience with production-level quantization on real models is essential. Proficiency with at least two of the following runtimes is required: ONNX Runtime, TensorRT, CoreML, OpenVINO, QNN, or LiteRT. Demonstrating profiling fluency on at least one platform is also a must.

Essential Skills

ML SystemsPyTorchONNX ExportONNX RuntimeTensorRTCoreMLOpenVINOQNNLiteRTProfiling

Good to Have

Custom Op Authoring

Highlights

  • Actively hiring

More Details

RolePerformance Engineer, On-Device Inference
IndustryAI / Machine Learning
DepartmentSoftware Development
Employment TypeFull Time, Hybrid (office + remote)

About the Company

Sarvam AI logo

Sarvam AI

AI / Machine Learning

Performance Engineer, On-Device Inference at Sarvam AI | SkillMX | SkillMX