Senior Software Engineer, Fabric Networking - GPU

NVIDIA

5+ yrs Bengaluru, Hyderabad, Pune, Gurugram Full Time Hybrid (office + remote)
NVIDIA logo
Posted : today
Actively hiring

Job description

Join NVIDIA, a pioneer in AI, HPC, and Visualization, and contribute to the future of computing. We are seeking experienced Senior Software Engineers to enhance our GPU Fabric Networking group. This role involves developing and maintaining system software essential for high-speed GPU communication, powering next-generation AI and deep learning platforms.

NVIDIA is at the forefront of technological innovation, with our GPUs driving advancements across diverse fields. As part of our team, you will play a crucial role in accelerating the next wave of AI breakthroughs. This is an opportunity to work with cutting-edge technology and shape the future of high-performance computing.

Responsibilities

Build, deploy, and maintain system software for efficient inter-GPU communication.

Contribute to the design and development of next-generation communication hardware and software for large-scale computing.

Drive platform bring-up, feature enablement, and end-to-end software validation and debugging for NVLink-based systems.

Resolve complex software, firmware, networking, and platform issues across various environments.

Enhance software quality, performance, reliability, and scalability within the GPU networking stack.

Collaborate with cross-functional teams to define and deliver robust software solutions.

Qualifications

A Bachelor's degree or equivalent experience in Computer Science, Computer Engineering, or a related field, or a Master's degree or equivalent experience is required.

Possess at least 5 years of professional software engineering experience.

Demonstrate excellent C/C++ programming, debugging, and problem-solving abilities.

Proficiency in shell scripting is essential; Python and Perl experience is a valuable asset.

Experience developing software that interacts with device drivers and exposes hardware functionality is necessary.

Exhibit a strong understanding of computer architecture, operating systems, Linux kernel internals, and systems programming.

Solid Linux development background and proficiency across multiple platforms, including Linux and Windows, are expected.

Experience developing multi-core, multi-process, and multi-threaded applications is required.

Strong fundamentals in networking, including TCP/IP, Ethernet, InfiniBand, RDMA/RoCE, routing, switching, and fabric performance analysis.

Familiarity with operating system virtualization technologies like KVM, QEMU, or Hyper-V is required.

Excellent written and verbal communication skills are necessary for effective collaboration.

Essential Skills

C++Shell ScriptingComputer ArchitectureOperating SystemsLinux KernelSystems ProgrammingLinuxMulti-core ProgrammingMulti-process ProgrammingMulti-threaded ProgrammingTCP/IPEthernetInfiniBandRDMA/RoCERoutingSwitchingFabric Performance AnalysisKVMQEMUHyper-V

Good to Have

PythonPerlNVIDIA GPU SystemsNVLinkNVSwitchCUDAAI/HPC ClustersPCIeDMAHigh-speed InterconnectsServer ManagementData Center OperationsCluster ProvisioningStatic Code AnalysisDynamic Code AnalysisFuzz TestingNegative TestingPerformance ValidationAI-assisted Development Tools

Highlights

  • Actively hiring

More Details

RoleSenior Software Engineer, Fabric Networking - GPU
IndustrySemiconductors, AI / Machine Learning
DepartmentSoftware Development
Employment TypeFull Time, Hybrid (office + remote)

About the Company

Nvidia logo

Nvidia

Semiconductors

Senior Software Engineer, Fabric Networking - GPU at NVIDIA | SkillMX | SkillMX