Senior Solution Architect, Cloud Infrastructure-DevOps
NVIDIA
NVIDIA
Join NVIDIA, a global leader in computer graphics, AI, and accelerated computing, as a Senior Cloud Infrastructure/DevOps Solutions Architect. You will be instrumental in building some of the world's largest and fastest AI/HPC systems. This dynamic, customer-focused role involves close collaboration with customers, partners, and internal teams to design and implement large-scale networking projects. Your expertise will span networking, system design, and automation, positioning you as the primary technical interface for clients.
Develop and maintain robust continuous integration and delivery pipelines. Create tools to automate the deployment and management of extensive infrastructure environments. Implement automated operational monitoring and alerting systems. Enable self-service consumption of resources through automation. Deploy comprehensive monitoring solutions for servers, networks, and storage. Troubleshoot issues across all layers, from bare metal and operating systems to software stacks and applications. Serve as a key technical resource, establishing and documenting standard methodologies for internal teams. Engage in Research & Development activities, including Proofs of Concept (POCs) and Proofs of Value (POVs) for future enhancements.
A Bachelor's or Master's degree in Computer Science, Data Science, Electrical/Computer Engineering, Physics, Mathematics, or a related engineering field, or equivalent practical experience, is required. A PhD is also welcome.
A minimum of 8 years of professional or research experience in networking fundamentals, TCP/IP stack, and data center architecture is essential.
More than 5 years of experience is needed in designing, implementing, and maintaining large-scale HPC/AI clusters with integrated monitoring, logging, and alerting. You must have experience managing Linux job/workload schedulers and orchestration tools.
Proficiency in HPC and AI solution technologies, including CPUs, GPUs, high-speed interconnects, and supporting software, is crucial.
Direct experience in designing, implementing, and managing cloud computing platforms such as AWS, Azure, or Google Cloud is required.
Strong knowledge of Windows and Linux networking (sockets, firewalld, iptables, Wireshark), system internals, ACLs, OS-level security, and common protocols (TCP, DHCP, DNS) is expected.
Experience with various storage solutions like Lustre, GPFS, ZFS, and XFS, along with familiarity with emerging storage technologies, is necessary.
Proficiency in Python programming and bash scripting, coupled with experience in automation and configuration management tools (Jenkins, Ansible, Puppet/Chef), is required.
Deep understanding of networking protocols like InfiniBand and Ethernet, and extensive experience with virtual systems (VMware, Hyper-V, KVM, Citrix), are essential.
Exceptional written, verbal, and listening communication skills in English are critical for this role.
Nvidia
Semiconductors