HPC Operations Engineer
NVIDIA
NVIDIA
Join NVIDIA, a leader in transforming computer graphics, PC gaming, and accelerated computing, as an HPC Operations Engineer. You will be instrumental in ensuring the seamless operation of our high-performance computing (HPC) environment, supporting semiconductor build and advanced engineering workflows. This is a unique opportunity to contribute to the future of computing within an innovative and supportive team.
Be part of a world-class team that powers cutting-edge semiconductor development and sophisticated engineering processes. Your role will be vital in maintaining the high availability and performance of our critical computing infrastructure.
Provide immediate support to HPC users for issues related to scheduling, compute resources, storage, and access.
Investigate and resolve job failures, scheduler errors, resource limitations, and performance bottlenecks, escalating when necessary.
Execute triage for infrastructure incidents, collecting necessary diagnostics and handing off complex issues to subject matter experts.
Actively monitor system health, job queues, node status, and service availability to guarantee stable daily operations.
Execute predefined operational procedures for system maintenance, software patching, and configuration updates.
Create and maintain comprehensive operational documentation, runbooks, and knowledge base articles for both internal teams and users.
Contribute to the enhancement of team processes by identifying recurring problems and suggesting practical workflow improvements.
A Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related discipline, or equivalent practical experience.
A minimum of 2 years of experience supporting production environments running on Linux.
Proficiency in Linux systems administration, with a strong understanding of RHEL/CentOS and/or Ubuntu.
Demonstrated ability to systematically troubleshoot technical challenges and identify critical customer concerns.
Prior experience engaging directly with users in a technical support or operations capacity.
Excellent written communication skills for producing clear documentation and procedural guides.
Proven capability to adhere to established processes with a keen attention to detail.
Bonus points for foundational scripting skills in Bash or Python for operational tasks, familiarity with workload schedulers (LSF, Slurm), understanding of network computing infrastructure (NFS, automounter, LDAP), and experience with HPC or large-scale compute environments, including EDA workloads.
Nvidia
Technology