Senior Systems Software Engineer - NV Cloud Functions
NVIDIA
NVIDIA
NVIDIA Cloud Functions (NVCF) is pioneering an open-source platform designed to seamlessly connect workloads with GPUs. This innovative system empowers teams to efficiently deploy, manage, and serve GPU-accelerated, containerized applications across global regions and clusters. By intelligently routing inference, streaming, and batch jobs through decentralized GPU clusters, NVCF ensures that endpoints can scale reliably, whether they are hosted on-premises or in the cloud. Join us as we push the boundaries of AI and redefine the future of computing.
As a Senior Systems Software Engineer, you will play a pivotal role in enhancing the performance, reliability, and scalability of a sophisticated system responsible for routing AI workloads across distributed GPU fleets. You will contribute to a polyglot, open-source platform, working on both control plane and edge deployments. This role is ideal for individuals with extensive experience in systems performance, distributed systems, and Kubernetes-based runtimes. You will design and deploy services using Java, Go, and Rust, contributing transparently to a public repository. Additionally, you'll automate and optimize build, test, integration, and release processes for cloud-native environments. Collaborating with diverse engineering teams across NVIDIA, you will ensure the platform's seamless integration with key technologies like KAI Scheduler, NVIDIA NIM, Grove, and Dynamo. You will also actively contribute to the open-source community by triaging issues, reviewing pull requests, and developing comprehensive documentation.
A Bachelor's or Master's Degree in Computer Science, or equivalent practical experience, is required. Candidates must possess a minimum of 3 years of hands-on software engineering experience. Essential qualifications include expert-level proficiency in a systems programming language such as Go, C, or Rust, coupled with a strong grasp of Data Structures, Algorithms, and Distributed Software Architecture. A robust understanding of container orchestration systems like Kubernetes, along with container technologies and automation experience in continuous integration frameworks (e.g., GitLab, ArgoCD), is crucial. Expertise in a scripting language like Bash or Python, and familiarity with the system internals of Unix/Unix-like kernels (especially Linux), are also necessary. A solid understanding of performance, security, and reliability principles in complex distributed systems is expected.
Nvidia
Technology