Principal Software Engineer, Distributed Systems Engineer - DGX Cloud

NVIDIA

15+ yrs Durham Full Time Remote
NVIDIA logo
Posted : 1 week ago
Actively hiring

Job description

NVIDIA is seeking experienced software engineers with Kubernetes expertise to contribute to scaling its AI Infrastructure. We are looking for innovative thinkers who can offer novel ideas and demonstrate strong execution capabilities. This role involves continuous learning, improvement, and evolution within a dynamic environment. You will play a key role in enhancing NVIDIA's ability to develop and deploy cutting-edge infrastructure solutions for diverse AI applications. If you are creative, passionate about Kubernetes and GPUs, and enjoy a fun work atmosphere, we encourage you to apply.

For two decades, NVIDIA has led visual computing, blending art and science. The invention of the GPU revolutionized modern visual computing, expanding into gaming, film production, product design, medical diagnosis, and scientific research. We are now entering the AI computing era, driven by a new computing paradigm: GPU deep learning.

Responsibilities

You will be an integral part of the DGX Cloud team, responsible for production systems that support large-scale, GPU-powered clusters for various AI workloads. This includes developing custom software for scheduling GPU resources within Kubernetes.

You will implement robust monitoring and health management systems to ensure industry-leading reliability, availability, and scalability of GPU assets. This involves leveraging multiple data streams, from GPU hardware diagnostics to cluster and network telemetry.

Collaborate with cross-functional teams at NVIDIA to guarantee the reliable and consistent performance of production AI clusters. Analyze system failures and enhance services through a structured incident management process.

Qualifications

We require direct experience in a software engineering role within a highly technical organization, with a proven track record of impactful work. Essential is software development experience with Kubernetes APIs and frameworks, going beyond simple cluster operation.

Candidates should be highly motivated with excellent communication skills, capable of working effectively with multi-functional teams, principles, and architects, and coordinating seamlessly across organizational boundaries and geographies.

We expect 15+ years of experience in similar roles and a background in large-scale production systems. Familiarity with common software engineering principles, tools, and techniques is necessary.

A Bachelor's degree in Computer Science, Engineering, Physics, Mathematics, or a comparable field, or equivalent experience, is required. Technical proficiency includes a systems programming language such as Go or Python, along with a solid understanding of data structures and algorithms.

Standout qualifications include technical expertise in managing and automating large-scale distributed systems independent of cloud providers. Advanced hands-on experience and deep understanding of cluster management systems like Kubernetes, Slurm, and Bright Cluster Manager are highly valued. Proven operational excellence in maintaining reliable and performant AI infrastructure will set candidates apart.

Essential Skills

KubernetesGoPython

Good to Have

SlurmBright Cluster Manager

Highlights

  • Actively hiring

More Details

RolePrincipal Software Engineer, Distributed Systems Engineer - DGX Cloud
Employment TypeFull Time, Remote

About the Company

Nvidia logo

Nvidia

IT Consulting