Senior Solutions Architect, Infiniband and Networking Ethernet - NVIS
NVIDIA
NVIDIA
Join NVIDIA's Infrastructure Specialist Team as a Senior Networking (ETH/IB) Solutions Architect. This role is pivotal in empowering academic and commercial groups to revolutionize deep learning, data analytics, and power data centers globally. You will be instrumental in building some of the world's largest and fastest AI/HPC systems.
We are seeking an individual with exceptional interpersonal skills, adept at working within a dynamic, customer-focused team. The position involves extensive interaction with customers, partners, and internal teams. Your primary focus will be analyzing, defining, and implementing large-scale networking projects, encompassing networking, system design, and automation, while acting as the primary customer interface.
Key responsibilities include developing AI/HPC infrastructure for both new and existing clientele. You will support the operational and reliability aspects of large-scale AI clusters, emphasizing performance at scale, real-time monitoring, logging, and alerting.
This role demands engagement throughout the entire service lifecycle, from initial conception and design through deployment, operation, and continuous refinement. Maintaining services post-launch by measuring and monitoring availability, latency, and overall system health is crucial.
You will also provide valuable feedback to internal teams by documenting issues, suggesting workarounds, and proposing improvements.
We require a BS/MS/PhD or equivalent experience in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields. A minimum of 5 years of professional experience in networking fundamentals, specifically within Ethernet or InfiniBand environments, is essential.
Demonstrated hands-on experience with network switch/router platforms such as Cumulus Linux, SONiC, IOS, JunosOS, and EOS is necessary. Solid working knowledge of Ethernet/InfiniBand/RDMA core principles is expected.
Proficiency in end-to-end IB/Eth cluster deployment, adapter configuration, and firmware maintenance is required, along with the ability to conduct professional performance benchmarking using mainstream RDMA testing tools. You must be capable of independently diagnosing and troubleshooting typical IB/Eth network anomalies, including link flapping, connection failures, and bandwidth/latency jitter issues.
Mastery of practical RDMA network optimization strategies, such as QP tuning, MTU configuration, and congestion control optimization, is a must. Hands-on experience with RDMA-accelerated business scenarios, including distributed storage and high-performance computing clusters, is also required.
Extensive experience delivering automated network provisioning solutions using tools like Ansible, Salt, and Python is critical. The ability to develop CI/CD pipelines for network operations is essential. Strong written, verbal, and listening skills in English are mandatory for this role.
Nvidia
IT Consulting