Senior Platform Engineer, Network Infrastructure
NVIDIA
NVIDIA
Join NVIDIA's Cloud Foundations Reliability team, a critical part of the Global Network Infrastructure organization. This team designs, deploys, and manages the Kubernetes-based platform essential for provisioning, monitoring, and operating NVIDIA's global network across diverse environments. We focus on the platform's architecture, lifecycle, and automation, ensuring standardized deployment and management of network platforms and services.
This role offers end-to-end ownership and production responsibility for the Kubernetes platform that powers GNI network systems. You will collaborate closely with engineering owners to resolve issues and drive changes, taking complex challenges from conception to successful production deployment. The position emphasizes bringing deep Kubernetes expertise and fostering consistent engineering practices across US and Bangalore teams.
Design, build, and operate a robust Kubernetes platform for GNI network automation, telemetry, and operations spanning data centers, colocation, and cloud environments.
Manage the complete lifecycle of GNI Kubernetes environments, including onboarding, upgrades, capacity planning, availability, and disaster recovery.
Develop high-quality software and automation for cluster provisioning, validation, upgrades, remediation, and secure multi-cluster delivery via GitOps.
Provide essential production support for network services hosted on the platform, partnering with service teams for seamless operation.
Diagnose and resolve complex issues within the Kubernetes platform and hosted services, addressing control-plane health, networking, storage, scheduling, and multi-cluster dependencies. Drive issues to verified resolution.
Establish comprehensive production-readiness and observability standards, defining health signals, capacity metrics, alerts, runbooks, and recovery procedures.
Participate actively in the CFR production on-call rotation, including off-hours and weekend coverage, leading incident response and driving corrective actions.
A Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience, is required.
Possess over 8 years of experience in building or operating production Kubernetes platforms, network infrastructure, or distributed systems.
Demonstrate deep expertise with Kubernetes at scale, covering cluster lifecycle management, upgrades, networking, storage, and recovery protocols.
Exhibit proficiency in at least one general-purpose programming language, such as Go or Python.
Showcase experience with GitOps, infrastructure as code (IaC), CI/CD pipelines, and automated production delivery methodologies.
Have a solid background in deploying and supporting network automation or telemetry services on Kubernetes.
Experience with production on-call rotations, incident response, root-cause analysis, and driving corrective actions to completion is essential.
Nvidia
IT Consulting