Senior Software Engineer, DGX Cloud Production Engineering

NVIDIA

8+ yrs Santa Clara Full Time Remote
NVIDIA logo
Posted : today
Actively hiring

Job description

NVIDIA DGX Cloud is at the forefront of building and managing vast GPU infrastructures for AI research and production. Join our production engineering team to develop the automation, tools, and operational systems essential for making GPU clusters highly reliable, scalable, and secure. This role focuses on Kubernetes-based infrastructure, GPU cluster operations, reliability, automation, GitOps, and ensuring seamless Day 2 operability across all DGX Cloud environments.

NVIDIA is a leader in AI, High-Performance Computing, and Visualization. The GPU, our core invention, powers modern computers and is central to our innovative products. We are a team of forward-thinking, dedicated individuals. If you are creative, driven, and self-motivated, we encourage you to apply.

Your base salary will be determined by your location, experience, and peer compensation. For Level 4, the range is $184,000 - $287,500 USD, and for Level 5, it's $224,000 - $356,500 USD. You will also receive equity and benefits. Applications will be accepted until September 27, 2026.

Responsibilities

Develop and manage automation for large-scale GPU clusters across NVIDIA Cloud Partners (NCP) and on-premises setups. Create tools and services for provisioning, validating, upgrading, monitoring, repairing, and managing the entire cluster lifecycle. Enhance Day 0, Day 1, and Day 2 workflows for efficient cluster setup, handover, and production operations. Minimize manual interventions in production through APIs, GitOps, automation, and agent-assisted processes. Participate in on-call rotations, incident response, debugging, and conduct thorough follow-up actions. Collaborate with platform, storage, networking, security, and workload teams to ensure infrastructure readiness for production.

Qualifications

A minimum of 8 years of experience in building or operating production infrastructure. Proficiency in programming languages such as Python, Go, or similar. Demonstrated experience with Linux, Kubernetes, containers, cloud infrastructure, or infrastructure automation. Proven ability to troubleshoot distributed systems in production environments. Excellent communication skills and the capacity to collaborate effectively across diverse teams. A Bachelor's or Master's degree in Computer Science, or equivalent practical experience.

Essential Skills

PythonGoLinuxKubernetesContainersCloud InfrastructureInfrastructure Automation

Good to Have

GPU InfrastructureKubernetes OperatorsGitOpsTerraformArgoCDFleet AutomationSLOsIncident ResponseObservabilityReliability PracticesBMaaSVMaaSManaged KubernetesMulti-cloud Infrastructure

Highlights

  • Actively hiring

More Details

RoleSenior Software Engineer, DGX Cloud Production Engineering
IndustryTechnology
DepartmentSoftware Development
Employment TypeFull Time, Remote

About the Company

Nvidia logo

Nvidia

Technology

Senior Software Engineer, DGX Cloud Production Engineering at NVIDIA | SkillMX | SkillMX