HPC Operations Engineer

NVIDIA

2+ yrs Bengaluru Full Time Hybrid (office + remote)
NVIDIA logo
Posted : today
Actively hiring

Job description

Join NVIDIA, a leader in transforming computer graphics, PC gaming, and accelerated computing, as an HPC Operations Engineer. You will be instrumental in ensuring the seamless operation of our high-performance computing (HPC) environment, supporting semiconductor build and advanced engineering workflows. This is a unique opportunity to contribute to the future of computing within an innovative and supportive team.

Be part of a world-class team that powers cutting-edge semiconductor development and sophisticated engineering processes. Your role will be vital in maintaining the high availability and performance of our critical computing infrastructure.

Responsibilities

Provide immediate support to HPC users for issues related to scheduling, compute resources, storage, and access.

Investigate and resolve job failures, scheduler errors, resource limitations, and performance bottlenecks, escalating when necessary.

Execute triage for infrastructure incidents, collecting necessary diagnostics and handing off complex issues to subject matter experts.

Actively monitor system health, job queues, node status, and service availability to guarantee stable daily operations.

Execute predefined operational procedures for system maintenance, software patching, and configuration updates.

Create and maintain comprehensive operational documentation, runbooks, and knowledge base articles for both internal teams and users.

Contribute to the enhancement of team processes by identifying recurring problems and suggesting practical workflow improvements.

Qualifications

A Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related discipline, or equivalent practical experience.

A minimum of 2 years of experience supporting production environments running on Linux.

Proficiency in Linux systems administration, with a strong understanding of RHEL/CentOS and/or Ubuntu.

Demonstrated ability to systematically troubleshoot technical challenges and identify critical customer concerns.

Prior experience engaging directly with users in a technical support or operations capacity.

Excellent written communication skills for producing clear documentation and procedural guides.

Proven capability to adhere to established processes with a keen attention to detail.

Bonus points for foundational scripting skills in Bash or Python for operational tasks, familiarity with workload schedulers (LSF, Slurm), understanding of network computing infrastructure (NFS, automounter, LDAP), and experience with HPC or large-scale compute environments, including EDA workloads.

Essential Skills

Linux Systems AdministrationTechnical SupportDocumentationTroubleshootingCommunication

Good to Have

Bash ScriptingPython ScriptingWorkload SchedulersNFSAutomounterLDAPHPCEDA Workloads

Highlights

  • Actively hiring

More Details

RoleHPC Operations Engineer
IndustryTechnology
DepartmentOperations
Employment TypeFull Time, Hybrid (office + remote)

About the Company

Nvidia logo

Nvidia

Technology

HPC Operations Engineer at NVIDIA | SkillMX | SkillMX