Senior Applied Research Engineer - Accelerator Programming Model and Compiler
NVIDIA
NVIDIA
Join a pioneering team at the forefront of accelerated computing. We are seeking a seasoned Applied Research Engineer to revolutionize the next generation of programming models for NVIDIA's Programmable Vision Accelerator (PVA). This role bridges cutting-edge research with robust systems software, focusing on creating intuitive interfaces for AI coding agents and developers.
Your work will shape how workloads are expressed, integrating high-level abstractions, domain-specific languages, and advanced compiler/runtime interfaces. You'll delve into public APIs, debug complex LLVM-based backends, and refine AI-generated code for optimal performance and correctness. This is an opportunity to influence the future of AI development on NVIDIA's versatile platforms.
Drive the evolution of the PVA programming model, enhancing ease of use for developers and AI agents in creating optimized algorithms.
Abstract hardware intricacies into declarative interfaces accessible to AI agents, improving compile times and diagnostic tools.
Define the architecture and feature set for the PVA SDK, runtime APIs, and programming model.
Develop and refine benchmarks for evaluating AI-generated PVA code, informing compiler optimizations and diagnostics.
Enhance the LLVM-based VPU compiler backend targeting VLIW/SIMD architectures.
Research and implement efficient integration models for PVA workloads within CUDA-based heterogeneous pipelines, covering execution, memory management, and synchronization.
Collaborate closely with internal and external stakeholders to foster adoption and refine PVA runtime APIs.
A strong foundation with a BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related discipline, or equivalent practical experience.
Over ten years of experience developing high-performance low-level systems software, accelerator software, embedded systems, or compiler/toolchain infrastructure.
Proven ability in creating high-level programming abstractions, DSLs, or compiler IRs that effectively mask hardware complexity while maintaining peak performance.
Substantial experience in compiler, debugger, linker, or toolchain development, with a preference for LLVM expertise.
Hands-on experience developing code with AI agents such as Claude Code, Cursor, and OpenAI, utilizing inference SDKs.
Background in integrating compilers and developer tools with AI coding agents or agentic development frameworks.
Proficiency in programming SIMD/VLIW processors.
Excellent C++ development skills, including deep-level debugging and performance analysis.
Experience with development environments like Linux or QNX.
Exceptional communication and interpersonal skills.
Nvidia
Semiconductors