Principal Software Engineer, Distributed Systems Engineer - DGX Cloud
NVIDIA
NVIDIA
NVIDIA is seeking experienced software engineers with Kubernetes expertise to contribute to scaling its AI Infrastructure. We are looking for innovative thinkers who can offer novel ideas and demonstrate strong execution capabilities. This role involves continuous learning, improvement, and evolution within a dynamic environment. You will play a key role in enhancing NVIDIA's ability to develop and deploy cutting-edge infrastructure solutions for diverse AI applications. If you are creative, passionate about Kubernetes and GPUs, and enjoy a fun work atmosphere, we encourage you to apply.
For two decades, NVIDIA has led visual computing, blending art and science. The invention of the GPU revolutionized modern visual computing, expanding into gaming, film production, product design, medical diagnosis, and scientific research. We are now entering the AI computing era, driven by a new computing paradigm: GPU deep learning.
You will be an integral part of the DGX Cloud team, responsible for production systems that support large-scale, GPU-powered clusters for various AI workloads. This includes developing custom software for scheduling GPU resources within Kubernetes.
You will implement robust monitoring and health management systems to ensure industry-leading reliability, availability, and scalability of GPU assets. This involves leveraging multiple data streams, from GPU hardware diagnostics to cluster and network telemetry.
Collaborate with cross-functional teams at NVIDIA to guarantee the reliable and consistent performance of production AI clusters. Analyze system failures and enhance services through a structured incident management process.
We require direct experience in a software engineering role within a highly technical organization, with a proven track record of impactful work. Essential is software development experience with Kubernetes APIs and frameworks, going beyond simple cluster operation.
Candidates should be highly motivated with excellent communication skills, capable of working effectively with multi-functional teams, principles, and architects, and coordinating seamlessly across organizational boundaries and geographies.
We expect 15+ years of experience in similar roles and a background in large-scale production systems. Familiarity with common software engineering principles, tools, and techniques is necessary.
A Bachelor's degree in Computer Science, Engineering, Physics, Mathematics, or a comparable field, or equivalent experience, is required. Technical proficiency includes a systems programming language such as Go or Python, along with a solid understanding of data structures and algorithms.
Standout qualifications include technical expertise in managing and automating large-scale distributed systems independent of cloud providers. Advanced hands-on experience and deep understanding of cluster management systems like Kubernetes, Slurm, and Bright Cluster Manager are highly valued. Proven operational excellence in maintaining reliable and performant AI infrastructure will set candidates apart.
Nvidia
IT Consulting