Senior Applied Research Engineer - Accelerator Programming Model and Compiler

NVIDIA

10+ yrs Bengaluru Full Time Hybrid (office + remote)
NVIDIA logo
Posted : today
Actively hiring

Job description

Join a pioneering team at the forefront of accelerated computing. We are seeking a seasoned Applied Research Engineer to revolutionize the next generation of programming models for NVIDIA's Programmable Vision Accelerator (PVA). This role bridges cutting-edge research with robust systems software, focusing on creating intuitive interfaces for AI coding agents and developers.

Your work will shape how workloads are expressed, integrating high-level abstractions, domain-specific languages, and advanced compiler/runtime interfaces. You'll delve into public APIs, debug complex LLVM-based backends, and refine AI-generated code for optimal performance and correctness. This is an opportunity to influence the future of AI development on NVIDIA's versatile platforms.

Responsibilities

Drive the evolution of the PVA programming model, enhancing ease of use for developers and AI agents in creating optimized algorithms.

Abstract hardware intricacies into declarative interfaces accessible to AI agents, improving compile times and diagnostic tools.

Define the architecture and feature set for the PVA SDK, runtime APIs, and programming model.

Develop and refine benchmarks for evaluating AI-generated PVA code, informing compiler optimizations and diagnostics.

Enhance the LLVM-based VPU compiler backend targeting VLIW/SIMD architectures.

Research and implement efficient integration models for PVA workloads within CUDA-based heterogeneous pipelines, covering execution, memory management, and synchronization.

Collaborate closely with internal and external stakeholders to foster adoption and refine PVA runtime APIs.

Qualifications

A strong foundation with a BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related discipline, or equivalent practical experience.

Over ten years of experience developing high-performance low-level systems software, accelerator software, embedded systems, or compiler/toolchain infrastructure.

Proven ability in creating high-level programming abstractions, DSLs, or compiler IRs that effectively mask hardware complexity while maintaining peak performance.

Substantial experience in compiler, debugger, linker, or toolchain development, with a preference for LLVM expertise.

Hands-on experience developing code with AI agents such as Claude Code, Cursor, and OpenAI, utilizing inference SDKs.

Background in integrating compilers and developer tools with AI coding agents or agentic development frameworks.

Proficiency in programming SIMD/VLIW processors.

Excellent C++ development skills, including deep-level debugging and performance analysis.

Experience with development environments like Linux or QNX.

Exceptional communication and interpersonal skills.

Essential Skills

C++LLVMCompiler DevelopmentLow-level DebuggingPerformance ProfilingSIMD/VLIW ProcessorsLinuxQNX

Good to Have

CUDAOpenCLMLIRHalideTVMTritonGraph CompilersImage Processing DSLsISO 26262IEC 61508Agent HarnessesMulti-agent Workflows

Highlights

  • Actively hiring

More Details

RoleSenior Applied Research Engineer - Accelerator Programming Model and Compiler
IndustrySemiconductors
DepartmentSoftware Development
Employment TypeFull Time, Hybrid (office + remote)

About the Company

Nvidia logo

Nvidia

Semiconductors