Senior System Software Engineer, Software Defined Networking
NVIDIA
NVIDIA
Join NVIDIA's AI Cloud team as a Senior System Software Engineer specializing in Software Defined Networking (SDN). You will be instrumental in designing, developing, and managing high-performance, scalable SDN solutions for our AI Clouds, supporting demanding workloads like hyperscale multi-node training, inference, cloud gaming, and cloud functions.
This pivotal role encompasses the entire SDN stack lifecycle. You'll architect and implement cutting-edge control and data plane software, while also ensuring robust operational excellence through advanced reliability engineering, CI/CD practices, comprehensive observability, and effective incident response.
If you have a passion for building foundational cloud infrastructure and a deep understanding of modern networking, we encourage you to apply and shape the future of AI computing.
Architect and develop next-generation multi-tenant cloud SDN control and data plane software, including OVS, OVN, and OpenFlow. Build Infrastructure-as-a-Service virtual network orchestration and services using gRPC and REST to guarantee tenant workload security and performance SLAs for BMaaS, VMaaS, and Kubernetes. Contribute significantly to upstream OVN-Kubernetes and related open-source projects. Develop advanced software for network observability, encompassing monitoring, telemetry, intelligent metering, and performance analysis. Operate and support OVS-OVN based SDN solutions within NVIDIA's expansive AI Cloud environments. Ensure end-to-end observability for the SDN stack by developing and maintaining monitoring, alerting, distributed tracing, and dashboarding tools for real-time network health and performance insights. Design, enhance, and maintain CI/CD pipelines (GitLab) for Linux host networking, OVS, OVN, and Kubernetes CNIs. Implement GitOps or similar approaches for secure, seamless cloud infrastructure integration and drive reliability via incident management, resource monitoring, and performance tuning. Collaborate effectively with SRE, DevOps, and network engineering teams to ensure production readiness and enhance operational tooling.
A Bachelor's or Master's degree in Computer Science or a related technical field, or equivalent practical experience is required. Possess 5+ years of demonstrated experience in software development for large-scale distributed systems. Exhibit expert-level knowledge of OVN, OVS, OpenFlow, and contemporary network protocols. Demonstrate strong programming proficiency in C and Go, coupled with advanced scripting skills in Bash and Python. Showcase deep understanding of Kubernetes and practical experience deploying and supporting Container Network Interfaces (CNIs), specifically OVN-Kubernetes. Have hands-on experience with Infrastructure-as-Code and deployment tools such as Ansible, Terraform, ArgoCD, and Flux. Possess experience in designing and operating complex, multi-stage CI/CD pipelines. Provide hands-on experience developing secure, high-performance services utilizing gRPC and REST, incorporating TLS and robust authentication mechanisms. Exhibit strong knowledge of datacenter routing, switching, and Linux host/VM networking fundamentals.
Nvidia
Technology