Senior Software Engineer, DGX Cloud Production Engineering
NVIDIA
NVIDIA
NVIDIA DGX Cloud is at the forefront of building and managing extensive GPU infrastructure tailored for AI research and production workloads. We are seeking experienced Senior Software Engineers to contribute to the development of robust automation, essential tooling, and sophisticated operational systems that ensure the reliability, scalability, and safety of our GPU clusters.
This role is integral to a production engineering team specializing in Kubernetes-based infrastructure, GPU cluster operations, reliability engineering, automation, GitOps practices, and ensuring seamless Day 2 operability across all DGX Cloud environments. Join us to shape the future of AI infrastructure.
Develop and manage automated systems for large-scale GPU clusters, supporting both NVIDIA Cloud Partners (NCP) and on-premise deployments.
Create and maintain tools and services for cluster provisioning, validation, upgrades, monitoring, repair, and full lifecycle management.
Enhance Day 0, Day 1, and Day 2 workflows to streamline cluster setup, handoff, and ongoing production operations.
Minimize manual interventions in production through advanced APIs, GitOps, automation, and agent-assisted solutions.
Participate actively in on-call rotations, incident response efforts, debugging complex issues, and executing thorough follow-up actions.
Possess a minimum of 8 years of experience in building or operating production infrastructure.
Demonstrate strong proficiency in programming languages such as Python or Go.
Showcase hands-on experience with Linux, Kubernetes, container technologies, cloud environments, or infrastructure automation tools.
Exhibit a proven ability to troubleshoot distributed systems in live production settings.
Maintain clear communication skills and a collaborative spirit to work effectively across diverse teams. A Bachelor's or Master's degree in Computer Science or equivalent practical experience is required.
Nvidia
IT Consulting