Senior Software Engineer, Capacity Management - DGX Cloud
NVIDIA
NVIDIA
Join NVIDIA DGX Cloud, a premier platform for enterprise-level AI development and deployment. As the demand for accelerated computing surges, mastering capacity management is crucial for delivering exceptional customer experiences and optimizing our valuable GPU infrastructure.
This Senior Software Engineer role focuses on architecting and developing sophisticated systems that seamlessly integrate customer demand, resource availability, reservations, allocation, and utilization across the DGX Cloud ecosystem. You will collaborate closely with diverse teams, including engineering, product, operations, finance, and business, to translate intricate capacity data and operational workflows into robust software solutions and automated decision-making processes.
NVIDIA is at the forefront of innovation in AI, High-Performance Computing, and Visualization, powered by our revolutionary GPU technology. We are seeking talented individuals to help shape the future of AI and contribute to a dynamic and inclusive work environment. NVIDIA is recognized as a top desirable employer, known for its forward-thinking culture and dedicated workforce.
Key responsibilities include designing and implementing distributed services and data pipelines for comprehensive capacity planning, allocation, reservations, and utilization.
You will develop a unified model to represent available, committed, and forecasted GPU capacity across various cloud providers, regions, clusters, and products. A significant focus will be on automating capacity management workflows that currently rely on manual processes.
Building APIs, tools, and integrations to empower other DGX Cloud systems and teams to make informed, capacity-aware decisions is essential. You will also enhance forecasting, scenario planning, and operational visibility by integrating demand signals with infrastructure supply data.
Establishing robust monitoring, data quality controls, and service-level indicators for capacity systems is a core duty. Leading technical design reviews, setting engineering standards, and mentoring fellow engineers will contribute to team growth. You will also be tasked with diagnosing complex production issues and driving improvements in the reliability, performance, and scalability of capacity management services.
A Bachelor of Science degree or equivalent experience in Computer Science, Computer Engineering, or a related technical discipline is required.
We are looking for at least 12 years of software engineering experience, specifically in building production-grade systems. Proficiency in programming languages such as Python, Go, or Java is essential.
Demonstrated experience in designing distributed systems, backend services, APIs, and data processing pipelines is critical. Familiarity with cloud infrastructure, Kubernetes, compute platforms, or large-scale resource management systems is highly valued.
A strong understanding of data modeling, system integration, observability, and production operations is necessary. The ability to translate ambiguous business and operational needs into clear technical designs is key. Excellent communication skills and experience working collaboratively across engineering and non-engineering teams are expected.
Nvidia
Software Development