Technical Support Engineer – Slurm

NVIDIA

5+ yrs India Full Time Remote
NVIDIA logo
Posted : yesterday
Actively hiring

Job description

Join NVIDIA and contribute to the future of AI and accelerated computing. This role focuses on supporting Slurm, a critical workload manager for demanding AI and high-performance computing environments. You'll be part of a specialized team, managing complex support cases to ensure customers operate efficient and scalable clusters.

This position demands extensive production experience with Slurm and the ability to troubleshoot issues across the scheduler and its supporting infrastructure, including Linux, networking, storage, authentication, databases, and GPU systems.

Responsibilities

Own and resolve Slurm support cases for customers running production AI and HPC clusters, from initial investigation to final resolution.

Diagnose intricate problems related to Slurm daemons (slurmctld, slurmd, slurmdbd), job scheduling, node management, resource allocation, accounting, authentication, and high availability.

Address Slurm configuration and policy complexities, including partitions, reservations, priorities, fair-share, quality of service, backfill, preemption, GRES/TRES, cgroups, and job constraints.

Investigate performance, reliability, and scalability challenges using logs, diagnostic data, configuration analysis, reproductions, and source-level debugging when necessary.

Isolate issues across Slurm and its dependencies, such as Linux, MUNGE, databases, networking, parallel storage, containers, GPUs, and cluster management systems.

Provide expert guidance to customers on Slurm configuration, upgrades, operational best practices, resource management, and incident recovery.

Collaborate with engineering teams by creating clear technical descriptions, reproducible test cases, and well-substantiated defect reports.

Develop comprehensive guides, knowledge-base articles, diagnostic tools, and internal training materials to enhance Slurm expertise within the support organization.

Qualifications

A Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience, is required.

Possess 5+ years of hands-on experience administering and supporting Slurm in production HPC or AI environments, including experience with business-critical outage incidents.

Demonstrate an expert-level understanding of Slurm architecture, daemons, configuration, scheduling behavior, accounting, resource management, and common failure modes.

Exhibit the capacity to independently identify sophisticated Slurm incidents and guide them toward a technically sound resolution.

Showcase in-depth Linux system administration and problem-solving skills, including proficiency with systemd, cgroups, authentication mechanisms, networking, and database-backed services.

Have experience operating Slurm across multi-user clusters with intricate scheduling policies and diverse compute resources.

Strong analytical and research abilities are essential, with a demonstrated skill in differentiating Slurm defects from configuration, integration, infrastructure, or workload-related problems.

Excellent written and verbal communication skills are mandatory, enabling clear explanations of detailed technical findings and actionable recommendations.

Essential Skills

SlurmLinux System AdministrationNetworkingDatabasesHigh-Performance Computing (HPC)Artificial Intelligence (AI)TroubleshootingProblem SolvingCustomer SupportTechnical DocumentationCommunication SkillsSystemdCgroupsAuthenticationParallel StorageContainerization

Good to Have

Slurm Source CodeSPANKLuaPyxisEnrootApptainerSingularityNVIDIA Base Command ManagerBright Cluster Manager

Highlights

  • Actively hiring

More Details

RoleTechnical Support Engineer – Slurm
DepartmentSales
Employment TypeFull Time, Remote

About the Company

Nvidia logo

Nvidia

IT Consulting

Technical Support Engineer – Slurm at NVIDIA | SkillMX | SkillMX