Senior Solutions Architect, Networking and Compute Infrastructure

NVIDIA

5+ yrs Gurugram Full Time Hybrid (office + remote)
NVIDIA logo
Posted : today
Actively hiring

Job description

Join NVIDIA's Infrastructure Specialist Team as a Senior Solutions Architect, focusing on Networking and Compute Infrastructure. This role involves working with leading academic and commercial groups globally, leveraging NVIDIA products to advance deep learning, data analytics, and power data centers. Be part of building some of the world's largest and fastest AI/HPC systems. We seek a dynamic, customer-focused individual with excellent interpersonal skills to analyze, define, and implement large-scale networking projects. Your expertise will span networking, system design, and automation, serving as a key customer interface.

Responsibilities

Key responsibilities include architecting and building AI/HPC infrastructure for both new and existing customers. You will support the operational and reliability aspects of large-scale AI clusters, emphasizing performance at scale, real-time monitoring, logging, and alerting. The role involves managing the full service lifecycle from design and deployment to operation and refinement. Developing automation tooling for large-scale infrastructure, operational monitoring, and self-service resource consumption is crucial. You will also deploy monitoring solutions for servers, networks, and storage, and perform in-depth troubleshooting from bare metal to application level. Acting as a technical resource, you will develop and document standard methodologies for internal teams and customers, and engage in Proof of Concepts (POCs) for future enhancements.

Qualifications

We require a Bachelor's or Master's degree (or equivalent experience) in Computer Science, Data Science, Electrical/Computer Engineering, Physics, Mathematics, or related engineering fields, with a minimum of 5 years of experience in networking fundamentals, TCP/IP stack, and data center compute architecture. Essential knowledge includes HPC, AI, EVPN, BGP, OSPF, VXLAN protocols, and deep understanding of data center architecture fundamentals like compute, storage (PFS), InfiniBand, Ethernet, and NVLink. Experience with HPC performance benchmarks, cluster health checks, and profiling tools is necessary. Proficiency in Python, bash scripting, and advanced Linux knowledge is a must. Extensive experience delivering automated network provisioning and comfort with tools like Jenkins, Ansible, Puppet, and Chef are expected. Solid understanding of Ethernet/InfiniBand/RDMA core principles and excellent customer-facing communication skills are vital for engaging with customers, partners, and cross-functional teams across India. Willingness to travel is also required.

Essential Skills

NetworkingTCP/IPData Center ArchitectureLinux System AdministrationPythonBash ScriptingAutomationConfiguration ManagementEthernetInfiniBandRDMAHPCAIEVPNBGPOSPFVXLANNVLinkCustomer Facing Skills

Good to Have

KubernetesContainerizationMicroservicesRoCELinux CertificationsNetworking CertificationsNVIDIA CertificationsObservability

Highlights

  • Actively hiring

More Details

RoleSenior Solutions Architect, Networking and Compute Infrastructure
IndustryTechnology
DepartmentSolutions Architect, Infrastructure Engineer
Employment TypeFull Time, Hybrid (office + remote)

About the Company

Nvidia logo

Nvidia

Technology

Senior Solutions Architect, Networking and Compute Infrastructure at NVIDIA | SkillMX | SkillMX