Senior Software Engineer, Network Visibility Platform - DGX Cloud
NVIDIA
NVIDIA
Lead the development of NVIDIA's Global Network Visibility (GNV) platform, transforming network data into actionable insights. This role is critical for enhancing network health understanding, accelerating incident resolution, and confidently deploying AI, storage, backbone, and edge infrastructure. You will drive the platform's architectural vision, ensuring its credibility, widespread adoption, and trustworthiness.
Your contributions will define authoritative network data standards, scalable measurement methodologies, and impactful customer-facing visibility products. This includes establishing robust interfaces, guiding cross-domain technical decisions, and making strategic build-versus-buy determinations to deliver a cohesive and reliable platform experience.
Architect a comprehensive platform integrating network topology, configuration, telemetry, active measurements, change intelligence, and self-service functionalities.
Develop an authoritative network model that unifies physical and logical topologies, configurations, and service dependencies, actively identifying configuration errors during cluster bring-up.
Formulate a telemetry and measurement strategy encompassing reachability tests, streaming signals, traffic data, and event tracking.
Deliver intuitive products for health monitoring, path analysis, blast radius assessment, and self-service exploration, empowering users to resolve network issues independently.
Establish clear interfaces, data contracts, quality benchmarks, and operational guidelines to foster seamless team contributions and maintain a unified user experience.
Drive cross-functional roadmaps, expertly balancing trade-offs between development speed, cost, reliability, and maintainability, while mentoring engineers on critical design aspects.
Possess a Bachelor’s degree or equivalent practical experience, coupled with over 10 years of relevant industry experience in software or infrastructure platforms.
Demonstrated success in architecting and operating large-scale software or infrastructure platforms, with advanced proficiency in systems programming languages like Python or Go.
Exhibit deep expertise in at least one of the following areas: network topology and data sources, telemetry and active measurement, or observability products, complemented by a solid understanding of the others.
Strong foundation in data center and backbone networking, covering physical connectivity, routing protocols, network overlays, service dependencies, and common failure modes.
Practical experience with observability toolchains such as Prometheus, Grafana, and OpenTelemetry, including active monitoring, synthetic testing, and containerized deployments on Kubernetes.
Experience with large-scale data systems, including graph, relational, time-series, streaming, API, or event-driven architectures, capable of data reconciliation from multiple sources.
Proven technical leadership across teams, demonstrating strong product acumen, operational judgment, and the ability to make informed trade-offs regarding performance, resiliency, usability, and cost.
Nvidia
Semiconductors