Senior Systems Software Engineer - NV Cloud Functions
NVIDIA
NVIDIA
Join NVIDIA's Cloud Functions team as a motivated, product-minded AI/ML Engineer specializing in large-scale AI platform engineering. Our team develops and maintains a serverless deployment platform designed for AI applications. This platform facilitates and scales AI inferencing workloads by orchestrating distributed tasks across GPU-enabled, cloud-agnostic Kubernetes clusters.
You will collaborate with skilled engineers dedicated to rapid innovation, aiming to deliver the best possible product for both external clients and internal NVIDIA teams. This role is at the forefront of defining cloud engineering paradigms for AI at scale.
Become a key subject matter expert by thoroughly understanding user challenges and constraints, translating them into actionable product requirements and solutions to speed up AI model deployment and inference on the NVCF platform.
Lead the implementation of critical features, conducting user acceptance, load, and performance testing. Focus on enhancing customer experience, optimizing performance, and ensuring platform reliability.
Mentor and collaborate with other engineering teams developing products on our platform. Share best practices for high-performance, large-scale AI/ML workloads, including ML reliability engineering.
Develop reference architectures to guide customer use cases, incorporating the latest NVCF features and AI technologies.
Proactively manage customer issues to resolution, providing timely alerts for emerging issues and risks.
Evaluate emerging technologies and tools within the evolving AI-at-scale landscape to maintain a competitive product and a forward-looking roadmap.
We seek candidates with a Master's, PhD, or equivalent experience in Computer Science, Artificial Intelligence, Applied Math, or a related discipline.
A minimum of 2 years of professional experience with Python, Rust, Golang, Linux, or Bash is required.
Demonstrated experience in Deep Learning and Machine Learning is essential, with expertise in AI/DL frameworks and inferencing software like SGLang, vLLM, TensorRT-LLM, or Dynamo.
Knowledge of CPU and GPU architecture is necessary.
Exceptional interpersonal skills are vital, including the ability to articulate complex technical concepts to non-expert audiences.
Experience in designing, implementing, and releasing AI/ML products to the market is expected. A flexible approach to technology and a comprehensive understanding of the software development lifecycle are key.
Nvidia
Technology