Senior AI Infrastructure Software Engineer - DGX Cloud
NVIDIA
NVIDIA
Join NVIDIA's DGX Cloud Lepton Team and contribute to a premier cloud product that empowers AI researchers and developers worldwide. Our focus is on building advanced AI/ML platforms designed to boost productivity, enhance the efficiency and resilience of AI workloads, and develop globally scalable AI infrastructure services. We are seeking a skilled AI Infrastructure Software Engineer to enhance our team.
You will play a crucial role in architecting, developing, and maintaining robust AI platforms essential for large-scale AI training, inferencing, fine-tuning, and production-level Agentic AI.
This senior role offers the chance to work with cutting-edge technologies shaping the future of AI within a collaborative and supportive team environment that champions learning and professional growth. You'll find the autonomy to tackle significant projects with the necessary guidance and mentorship, contributing to a culture that values blameless postmortems, continuous improvement, and strategic risk-taking. If you're looking for a stimulating and impactful career, we encourage you to apply.
Develop sophisticated platforms and tools specifically for large-scale AI, LLM, and GenAI infrastructure.
Create and refine tools to maximize the efficiency and resilience of AI/ML workloads.
Conduct thorough root cause analysis, troubleshooting failures from the application layer all the way to the hardware.
Advance the infrastructure and products that form the backbone of NVIDIA's AI platforms.
Collaborate on the design and implementation of APIs for seamless integration with NVIDIA's platform resiliency stacks.
Define and track key reliability metrics to drive improvements in system and service dependability.
A minimum of 8 years of experience in developing software infrastructure for large-scale AI systems is required.
A Bachelor's degree in Computer Science or a related technical field, or equivalent practical experience, is necessary.
Possess strong debugging capabilities and demonstrated experience in analyzing and triaging AI applications across all layers, from application to hardware.
Showcase a proven history of building and scaling large-scale distributed systems effectively.
Familiarity with AI training, inferencing, and data infrastructure services is essential.
Experience with Kubernetes and managing large-scale observability platforms for monitoring and logging (e.g., ELK, Prometheus, Loki) is expected.
Proficiency in programming languages such as Golang, Python, C/C++, and scripting languages is a must.
Excellent communication and collaboration skills, coupled with a mindset of diversity, intellectual curiosity, problem-solving, and openness, are critical for success.
Nvidia
Technology