Senior HPC Platform Architect

NVIDIA

5+ yrs Bengaluru Full Time Work from office
NVIDIA logo
Posted : today
Actively hiring

Job description

Join NVIDIA's cutting-edge HPC Infrastructure team as a Senior HPC Platform Architect. This role is crucial for designing, evaluating, and optimizing the compute infrastructure that powers NVIDIA's next-generation silicon design and AI workloads. You will be instrumental in architecting new data center clusters globally, acting as the primary reviewer and performance advocate.

This is an exciting opportunity to shape the future of high-performance computing infrastructure at a leading technology company.

Responsibilities

Your core responsibilities will involve conducting thorough data center architecture reviews for new HPC clusters, scrutinizing decisions regarding compute, storage, networking, and cooling. You will represent the BDC in cluster build meetings across diverse programs, ensuring architectural alignment. Analyzing and validating cluster design choices, including storage-to-compute distance, latency, rack layout, and multi-site topology, will be key to surfacing risks and recommending tradeoffs.

Additionally, you will lead performance benchmarking and profiling efforts to pinpoint bottlenecks early. Driving infrastructure optimization across multiple layers—from scheduler tuning to OS/kernel-level adjustments—is essential. Collaboration with platform and operations teams for cluster health and capacity planning, alongside evaluating new hardware with vendors, will be integral to your role. You will also focus on continuously enhancing infrastructure observability and documentation.

Qualifications

We are seeking candidates with a B.E./B.Tech or M.Tech/M.S. degree and at least 5 years of experience in HPC infrastructure, data center architecture, systems engineering, or a senior SRE/platform engineering role. A deep understanding of data center fundamentals, including compute (CPU/GPU servers), storage (parallel file systems, NVMe, tiered storage), and high-speed networking (InfiniBand, Ethernet, NVLink), is required.

Proven experience in evaluating and challenging infrastructure design decisions is crucial. Expertise in OS and kernel-level performance tuning, hands-on administration of large-scale Linux HPC clusters with workload managers like LSF/Slurm, and strong Linux/Unix system administration skills with scripting proficiency (Python, Bash, Perl) are essential. Experience with HPC performance benchmarks, cluster health checks, and profiling tools is also necessary. Excellent problem-solving and communication skills are vital for synthesizing complex tradeoffs and presenting clear recommendations to leadership.

Essential Skills

HPCData Center ArchitectureSystems EngineeringLinuxPythonBashPerlInfiniBandEthernetNVLinkLSFSlurmPerformance BenchmarkingSystem AdministrationTroubleshootingCommunication

Highlights

  • Actively hiring

More Details

RoleSenior HPC Platform Architect
IndustrySemiconductors, AI / Machine Learning
DepartmentEngineering, Information Technology, Computer Science
Employment TypeFull Time, Work from office

About the Company

Nvidia logo

Nvidia

Semiconductors

Senior HPC Platform Architect at NVIDIA | SkillMX | SkillMX