Senior Software Engineer, Distributed Systems Engineer - DGX Cloud

NVIDIA

5+ yrs Santa Clara Full Time Remote
NVIDIA logo
Posted : today
Actively hiring

Job description

NVIDIA is seeking seasoned software engineers with extensive Kubernetes experience to bolster our AI Infrastructure scaling efforts. We are looking for innovative thinkers who bring fresh ideas and a drive for execution. This role offers continuous opportunities for learning, improvement, and evolution, contributing to NVIDIA's leadership in building and deploying cutting-edge infrastructure solutions for diverse AI applications. If you're a creative problem-solver passionate about Kubernetes and GPUs, and enjoy a dynamic work environment, we encourage you to apply.

For two decades, NVIDIA has led visual computing, revolutionizing computer graphics. The invention of the GPU propelled advancements across gaming, film production, product design, medical diagnosis, and scientific research. Today, we are at the forefront of the AI computing era, powered by GPU deep learning.

Responsibilities

As a key member of the DGX Cloud team, you will manage production systems for large-scale GPU clusters supporting various AI workloads. Your responsibilities include developing custom software for Kubernetes-based GPU resource scheduling.

You will implement robust monitoring and health management solutions to ensure industry-leading reliability, availability, and scalability of GPU assets, leveraging multiple data streams from hardware diagnostics to cluster telemetry.

Collaboration with cross-functional NVIDIA teams is essential to guarantee the consistent, reliable, and high-performance operation of production AI clusters. You will analyze system failures and implement service improvements through a well-defined incident management process.

Qualifications

A proven track record in a software engineering role within a highly technical setting, demonstrating significant impact from your contributions, is required. Experience developing with Kubernetes APIs and frameworks, beyond mere cluster operation, is essential.

We seek highly motivated individuals with strong communication skills, capable of effective collaboration with multi-functional teams, principles, and architects, seamlessly coordinating across organizational boundaries and geographies.

Candidates should possess at least 5 years of relevant experience, with a focus on large-scale production systems. Familiarity with common software engineering principles, tools, and techniques is expected. A Bachelor's degree in Computer Science, Engineering, Physics, Mathematics, or a comparable degree or equivalent experience is necessary.

Essential technical knowledge includes proficiency in a systems programming language like Go or Python, coupled with a solid understanding of data structures and algorithms.

Essential Skills

KubernetesSoftware EngineeringGoPythonData StructuresAlgorithmsDistributed SystemsCluster ManagementAI Infrastructure

Good to Have

SlurmBright Cluster Manager

Highlights

  • Actively hiring

More Details

RoleSenior Software Engineer, Distributed Systems Engineer - DGX Cloud
IndustryTechnology
DepartmentSoftware Development
Employment TypeFull Time, Remote

About the Company

Nvidia logo

Nvidia

Technology

Senior Software Engineer, Distributed Systems Engineer - DGX Cloud at NVIDIA | SkillMX | SkillMX