Senior Applied Research Engineer, Accelerator Algorithms
NVIDIA
NVIDIA
Join a leading accelerated computing company and help shape the future of autonomous machines. This role focuses on applied research and architecture for algorithms on NVIDIA’s Programmable Vision Accelerator (PVA), a power-efficient VLIW/SIMD processor. You will analyze real-world AV and physical AI workloads, optimize algorithms for the PVA, and influence future hardware and software development. This is an opportunity to work on cutting-edge technology that powers innovations in self-driving cars, robotics, and AI.
As a Senior Applied Research Engineer, you will dive deep into complex AI workloads, prototype accelerator-friendly algorithms, and collaborate with research, compiler, and hardware teams. Your insights will directly impact the PVA's architecture, programming model, and software stack, driving performance, power, and latency benefits in final products.
Conduct applied research to characterize real-world AI and autonomous vehicle workloads, identifying optimal algorithms for the PVA. Develop and refine PVA algorithms, focusing on instruction-level parallelism, memory optimization, and efficient data movement. Translate workload and algorithm findings into concrete requirements for PVA hardware architecture, compilers, and SDKs. Build and utilize prototypes, benchmarks, and performance models to assess algorithm efficiency across current and future hardware. Employ AI-assisted development tools to accelerate workload analysis and algorithm creation. Collaborate with internal teams and external customers to integrate PVA-accelerated algorithms into production systems. Communicate technical findings and research results effectively across engineering teams.
Requires a BS/MS or PhD in a relevant technical field, or equivalent practical experience. Minimum of 12 years of experience in applied research, accelerator algorithm design, computer architecture, or HPC. Proficiency in DSP, SIMD, VLIW, fixed-point arithmetic, memory hierarchies, and low-level performance tuning. Experience with HW/SW co-design, workload analysis, performance modeling, and bottleneck identification. Familiarity with physical AI workloads such as VLA models, multimodal perception, and robotics. Strong programming skills in C++, Python, or CUDA. Excellent communication abilities for clearly explaining complex technical trade-offs. Demonstrated ownership, strong technical judgment, and ability to navigate ambiguity.
Nvidia
Semiconductors