Technical Support Engineer – Slurm
NVIDIA
NVIDIA
Join NVIDIA and contribute to the future of AI and accelerated computing. This role focuses on supporting Slurm, a critical workload manager for demanding AI and high-performance computing environments. You'll be part of a specialized team, managing complex support cases to ensure customers operate efficient and scalable clusters.
This position demands extensive production experience with Slurm and the ability to troubleshoot issues across the scheduler and its supporting infrastructure, including Linux, networking, storage, authentication, databases, and GPU systems.
Own and resolve Slurm support cases for customers running production AI and HPC clusters, from initial investigation to final resolution.
Diagnose intricate problems related to Slurm daemons (slurmctld, slurmd, slurmdbd), job scheduling, node management, resource allocation, accounting, authentication, and high availability.
Address Slurm configuration and policy complexities, including partitions, reservations, priorities, fair-share, quality of service, backfill, preemption, GRES/TRES, cgroups, and job constraints.
Investigate performance, reliability, and scalability challenges using logs, diagnostic data, configuration analysis, reproductions, and source-level debugging when necessary.
Isolate issues across Slurm and its dependencies, such as Linux, MUNGE, databases, networking, parallel storage, containers, GPUs, and cluster management systems.
Provide expert guidance to customers on Slurm configuration, upgrades, operational best practices, resource management, and incident recovery.
Collaborate with engineering teams by creating clear technical descriptions, reproducible test cases, and well-substantiated defect reports.
Develop comprehensive guides, knowledge-base articles, diagnostic tools, and internal training materials to enhance Slurm expertise within the support organization.
A Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience, is required.
Possess 5+ years of hands-on experience administering and supporting Slurm in production HPC or AI environments, including experience with business-critical outage incidents.
Demonstrate an expert-level understanding of Slurm architecture, daemons, configuration, scheduling behavior, accounting, resource management, and common failure modes.
Exhibit the capacity to independently identify sophisticated Slurm incidents and guide them toward a technically sound resolution.
Showcase in-depth Linux system administration and problem-solving skills, including proficiency with systemd, cgroups, authentication mechanisms, networking, and database-backed services.
Have experience operating Slurm across multi-user clusters with intricate scheduling policies and diverse compute resources.
Strong analytical and research abilities are essential, with a demonstrated skill in differentiating Slurm defects from configuration, integration, infrastructure, or workload-related problems.
Excellent written and verbal communication skills are mandatory, enabling clear explanations of detailed technical findings and actionable recommendations.
Nvidia
IT Consulting