Senior Performance Engineer, Intel Stack
Sarvam AI
Sarvam AI
Join Sarvam in building India's sovereign AI platform, focusing on research, models, infrastructure, and applications. We're dedicated to making AI work effectively for India and partner with leading enterprises. Be a key player in our mission to drive AI innovation and impact across the nation.
Take full ownership of Sarvam's Intel surface, covering Intel NPUs (OpenVINO, Meteor Lake, Lunar Lake, vPro AI PCs), Intel integrated and discrete GPUs (OpenVINO, ONNX Runtime), and x86/AMD64 CPU optimization. You will be the technical liaison with Intel's ecosystem and OpenVINO team, ensuring our edge models meet defined SLAs.
Key responsibilities include owning the OpenVINO build and quantization process, managing driver-version compatibility, and driving x86 CPU optimization for fallback paths using AVX-512 and AMX. You will also manage the Intel device-CI pool and detect regressions during OpenVINO upgrades.
We are seeking an experienced professional with over 5 years in ML deployment, including at least 2 years specifically on Intel inference stacks. Essential experience includes production-level OpenVINO usage, model conversion, accuracy validation post-quantization, and driver-version pinning. Familiarity with ONNX Runtime EPs and understanding when to utilize OpenVINO EP versus CPU EP is crucial.
Proficiency in x86 CPU profiling and optimization tools like VTune and perf is expected. While AVX-512 or AMX intrinsics are considered a strong plus, they are not strictly mandatory. Prior interaction with the OpenVINO team or Intel ecosystem partners, along with custom OpenVINO operator authoring, would be highly advantageous.
Sarvam AI
AI / Machine Learning