Senior System Software Engineer - Local AI
NVIDIA
NVIDIA
Join NVIDIA's LocalAI team as a Senior Systems Software Engineer, contributing to the forefront of on-device AI. This role focuses on architecting and developing efficient, high-performance local AI software for RTX and DGX-class systems. You'll be instrumental in reducing latency, enhancing real-time processing, and addressing privacy concerns by bringing AI capabilities closer to data generation points.
This is a unique opportunity to shape the future of client-side AI, driving innovation on resource-constrained platforms and playing a crucial role in the evolution of intelligent applications across various workflows.
Partner with NVIDIA's software, research, and product leadership, alongside external partners like Microsoft, to shape the AI ecosystem on RTX and DGX PCs.
Build and optimize the local AI inference stack for RTX, RTX Pro, and DGX GPUs, ensuring performance, stability, and scalability across diverse hardware.
Architect and develop modern inference runtimes using frameworks such as llama.cpp, vLLM, PyTorch, and TensorRT-RTX for various AI workloads, including LLM, vision-language, and TTS.
Perform end-to-end optimization of AI models and inference runtimes, applying techniques like quantization and pruning for efficient deployment on edge devices.
Lead system-level debugging, performance tuning, and analysis of performance-accuracy trade-offs. Establish engineering guidelines to accelerate bring-up and ensure production readiness.
Mentor junior engineers and review architecture proposals, fostering technical quality and consistency within the LocalAI organization.
A Bachelor's, Master's, or PhD in Computer Science, Software Engineering, Mathematics, or a related field, or equivalent experience, is required.
Possess 5+ years of experience in C++ programming, with a strong grasp of data structures, algorithms, and machine learning principles.
Demonstrate proven experience architecting and optimizing AI inference pipelines using frameworks like Llama.cpp, vLLM, PyTorch, and TensorRT.
Exhibit a deep understanding of inference backends and runtime internals, including scheduling, memory management, and hardware-aware optimization techniques.
Showcase the ability to establish technical direction and influence across multiple teams. Strong analytical and problem-solving skills are essential for managing priorities in a dynamic environment.
Excellent written and verbal communication skills are necessary for effective collaboration across engineering teams and management.
Nvidia
AI / Machine Learning