Senior Systems Software Engineer, Kubernetes Node Lifecycle - DGX Cloud
NVIDIA
NVIDIA
Join NVIDIA's DGX Cloud division, at the forefront of accelerated computing solutions for complex AI workloads. We are seeking a Senior Systems Software Engineer specializing in Kubernetes node engineering, OS image packaging, and cloud infrastructure. The ideal candidate will possess extensive hyperscaler-level expertise across the entire node lifecycle, from CAPI providers and bring-your-own-node onboarding to OS image build pipelines, packaging, and nodepool management. This role is crucial for maintaining cluster reliability at a frontier AI scale, managing the node layer within NVIDIA Kubernetes Engine (NKE) to support both internal researchers and NCPs.
This position offers an exciting opportunity to innovate and shape the future of AI infrastructure. You will play a key role in ensuring our solutions are robust, scalable, and secure, driving technological advancements that impact millions globally. If you're passionate about cutting-edge technology and making a difference, we encourage you to apply.
Key responsibilities include directing the development and enhancement of CAPI providers for NVIDIA Kubernetes Engine to ensure consistent and scalable node provisioning. You will develop and maintain bring-your-own-node workflows, enabling seamless integration of diverse NVIDIA hardware into NKE clusters. Coordinating OS image generation, packaging, deployment, and updates is essential, ensuring images are optimized for GPU workloads and meet enterprise-grade security standards.
Further duties involve developing and sustaining robust node image hardening pipelines, incorporating security benchmarks and automated vulnerability remediation. You'll also develop and maintain automated test suites for node images, ensuring accuracy across Kubernetes versions and hardware configurations. Handling large-scale nodepool lifecycles, including provisioning, upgrades, and replacements, is critical. Additionally, you will troubleshoot and resolve node-layer faults in production NKE clusters, and collaborate with upstream communities like Cluster API and Kubernetes to set node provisioning standards.
We are looking for candidates with at least 8 years of experience in systems software, cloud infrastructure, or Kubernetes node engineering. A Bachelor’s or Master’s degree in Engineering (Electrical, Computer Engineering, Computer Science) or equivalent practical experience is required.
Essential qualifications include deep expertise in Cluster API (CAPI), extensive experience with OS image build pipelines and node image packaging for Kubernetes, and practical experience with bring-your-own-node models. Strong understanding of kubelet configuration, node bootstrap, and Kubernetes node registration is necessary. Proficiency in Golang and/or Python, along with hands-on experience with major public cloud providers (GCP, AWS, Azure, OCI), is also required. Experience with node image security, including vulnerability scanning and patch automation, is a significant plus.
Nvidia
Technology