Senior Software Engineer, Network Visibility Platform - DGX Cloud

NVIDIA

10+ yrs Santa Clara Full Time Remote
NVIDIA logo
Posted : today
Actively hiring

Job description

Lead the development of NVIDIA's Global Network Visibility (GNV) platform, transforming network data into actionable insights. This role is critical for enhancing network health understanding, accelerating incident resolution, and confidently deploying AI, storage, backbone, and edge infrastructure. You will drive the platform's architectural vision, ensuring its credibility, widespread adoption, and trustworthiness.

Your contributions will define authoritative network data standards, scalable measurement methodologies, and impactful customer-facing visibility products. This includes establishing robust interfaces, guiding cross-domain technical decisions, and making strategic build-versus-buy determinations to deliver a cohesive and reliable platform experience.

Responsibilities

Architect a comprehensive platform integrating network topology, configuration, telemetry, active measurements, change intelligence, and self-service functionalities.

Develop an authoritative network model that unifies physical and logical topologies, configurations, and service dependencies, actively identifying configuration errors during cluster bring-up.

Formulate a telemetry and measurement strategy encompassing reachability tests, streaming signals, traffic data, and event tracking.

Deliver intuitive products for health monitoring, path analysis, blast radius assessment, and self-service exploration, empowering users to resolve network issues independently.

Establish clear interfaces, data contracts, quality benchmarks, and operational guidelines to foster seamless team contributions and maintain a unified user experience.

Drive cross-functional roadmaps, expertly balancing trade-offs between development speed, cost, reliability, and maintainability, while mentoring engineers on critical design aspects.

Qualifications

Possess a Bachelor’s degree or equivalent practical experience, coupled with over 10 years of relevant industry experience in software or infrastructure platforms.

Demonstrated success in architecting and operating large-scale software or infrastructure platforms, with advanced proficiency in systems programming languages like Python or Go.

Exhibit deep expertise in at least one of the following areas: network topology and data sources, telemetry and active measurement, or observability products, complemented by a solid understanding of the others.

Strong foundation in data center and backbone networking, covering physical connectivity, routing protocols, network overlays, service dependencies, and common failure modes.

Practical experience with observability toolchains such as Prometheus, Grafana, and OpenTelemetry, including active monitoring, synthetic testing, and containerized deployments on Kubernetes.

Experience with large-scale data systems, including graph, relational, time-series, streaming, API, or event-driven architectures, capable of data reconciliation from multiple sources.

Proven technical leadership across teams, demonstrating strong product acumen, operational judgment, and the ability to make informed trade-offs regarding performance, resiliency, usability, and cost.

Essential Skills

PythonGoNetwork TopologyTelemetryData Center NetworkingBackbone NetworkingRoutingOverlaysServicesKubernetesObservabilityPrometheusGrafanaOpenTelemetryGraph DatabasesTime-Series DatabasesAPIEvent-Driven Systems

Good to Have

AI NetworksHPC NetworksRDMARoCEInfiniBandNetBoxNautobotgRPCgNMITypeScriptReactSRENetwork Operations

Highlights

  • Actively hiring

More Details

RoleSenior Software Engineer, Network Visibility Platform - DGX Cloud
IndustrySemiconductors
DepartmentSoftware Development, DevOps / Cloud
Employment TypeFull Time, Remote

About the Company

Nvidia logo

Nvidia

Semiconductors

Senior Software Engineer, Network Visibility Platform - DGX Cloud at NVIDIA | SkillMX | SkillMX