Senior HPC Cluster Engineer - AI, ML
NVIDIA
NVIDIA
NVIDIA is a leader in accelerated computing, driving innovation in AI, ML, and high-performance computing. We are seeking a Senior AI/ML HPC Cluster Engineer for our Managed AI Superclusters (MARS) team. This role focuses on building and scaling infrastructure for advanced AI/ML systems, enabling researchers to develop next-generation solutions.
Join a team dedicated to designing and managing cutting-edge computing workloads. You will play a key role in shaping the future of AI infrastructure.
Lead systems administration and service delivery for our AI/HPC fleet, including system upgrades and incident response. Collaborate with global teams to enhance the AI and HPC research user experience. Oversee day-to-day operations of production AI/HPC clusters, ensuring system health and resource efficiency. Develop and refine our GPU-accelerated computing ecosystem through scalable automation. Build and maintain heterogeneous AI/ML clusters across on-premises and cloud environments. Foster strong relationships with users and internal teams to address evolving needs. Support researchers with workload execution, performance analysis, and optimization. Analyze and optimize cluster efficiency, minimizing fragmentation and GPU waste to meet SLAs. Conduct root cause analysis and implement proactive issue resolution. Lead SEV triage and postmortems for critical reliability incidents. Participate in on-call rotations for essential production GPU cluster support.
A Bachelor's degree in Computer Science, Electrical Engineering, or a related field, or equivalent practical experience, is required.
Minimum of 5 years' experience in designing and operating large-scale compute infrastructure.
Proficiency with advanced AI/HPC job schedulers like Slurm, K8s, PBS, RTDA, BCM, or LSF.
Strong administration skills for Centos/RHEL and/or Ubuntu Linux distributions.
Solid understanding of cluster configuration management tools (BCM, Terraform, Ansible, Puppet, Salt) and container technologies (Docker, Singularity, Podman, Shifter, Charliecloud).
Experience with Python programming and bash scripting.
Applied experience with AI/HPC workflows utilizing MPI.
Proven ability to analyze and tune performance for diverse AI/HPC workloads.
A passion for continuous learning and staying current with emerging technologies in HPC and AI/ML infrastructure.
Nvidia
Technology