Senior Software Engineer, DGX Cloud Production Engineering
NVIDIA
NVIDIA
NVIDIA DGX Cloud is at the forefront of building and managing vast GPU infrastructures for AI research and production. Join our production engineering team to develop the automation, tools, and operational systems essential for making GPU clusters highly reliable, scalable, and secure. This role focuses on Kubernetes-based infrastructure, GPU cluster operations, reliability, automation, GitOps, and ensuring seamless Day 2 operability across all DGX Cloud environments.
NVIDIA is a leader in AI, High-Performance Computing, and Visualization. The GPU, our core invention, powers modern computers and is central to our innovative products. We are a team of forward-thinking, dedicated individuals. If you are creative, driven, and self-motivated, we encourage you to apply.
Your base salary will be determined by your location, experience, and peer compensation. For Level 4, the range is $184,000 - $287,500 USD, and for Level 5, it's $224,000 - $356,500 USD. You will also receive equity and benefits. Applications will be accepted until September 27, 2026.
Develop and manage automation for large-scale GPU clusters across NVIDIA Cloud Partners (NCP) and on-premises setups. Create tools and services for provisioning, validating, upgrading, monitoring, repairing, and managing the entire cluster lifecycle. Enhance Day 0, Day 1, and Day 2 workflows for efficient cluster setup, handover, and production operations. Minimize manual interventions in production through APIs, GitOps, automation, and agent-assisted processes. Participate in on-call rotations, incident response, debugging, and conduct thorough follow-up actions. Collaborate with platform, storage, networking, security, and workload teams to ensure infrastructure readiness for production.
A minimum of 8 years of experience in building or operating production infrastructure. Proficiency in programming languages such as Python, Go, or similar. Demonstrated experience with Linux, Kubernetes, containers, cloud infrastructure, or infrastructure automation. Proven ability to troubleshoot distributed systems in production environments. Excellent communication skills and the capacity to collaborate effectively across diverse teams. A Bachelor's or Master's degree in Computer Science, or equivalent practical experience.
Nvidia
Technology