Senior Platform Engineer, Network Infrastructure

NVIDIA

8+ yrs Bengaluru Full Time Remote
NVIDIA logo
Posted : 2 weeks ago Actively Hiring

Job description

Join NVIDIA's Cloud Foundations Reliability team, a critical part of the Global Network Infrastructure organization. This team designs, deploys, and manages the Kubernetes-based platform essential for provisioning, monitoring, and operating NVIDIA's global network across diverse environments. We focus on the platform's architecture, lifecycle, and automation, ensuring standardized deployment and management of network platforms and services.

This role offers end-to-end ownership and production responsibility for the Kubernetes platform that powers GNI network systems. You will collaborate closely with engineering owners to resolve issues and drive changes, taking complex challenges from conception to successful production deployment. The position emphasizes bringing deep Kubernetes expertise and fostering consistent engineering practices across US and Bangalore teams.

Responsibilities

Design, build, and operate a robust Kubernetes platform for GNI network automation, telemetry, and operations spanning data centers, colocation, and cloud environments.

Manage the complete lifecycle of GNI Kubernetes environments, including onboarding, upgrades, capacity planning, availability, and disaster recovery.

Develop high-quality software and automation for cluster provisioning, validation, upgrades, remediation, and secure multi-cluster delivery via GitOps.

Provide essential production support for network services hosted on the platform, partnering with service teams for seamless operation.

Diagnose and resolve complex issues within the Kubernetes platform and hosted services, addressing control-plane health, networking, storage, scheduling, and multi-cluster dependencies. Drive issues to verified resolution.

Establish comprehensive production-readiness and observability standards, defining health signals, capacity metrics, alerts, runbooks, and recovery procedures.

Participate actively in the CFR production on-call rotation, including off-hours and weekend coverage, leading incident response and driving corrective actions.

Qualifications

A Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience, is required.

Possess over 8 years of experience in building or operating production Kubernetes platforms, network infrastructure, or distributed systems.

Demonstrate deep expertise with Kubernetes at scale, covering cluster lifecycle management, upgrades, networking, storage, and recovery protocols.

Exhibit proficiency in at least one general-purpose programming language, such as Go or Python.

Showcase experience with GitOps, infrastructure as code (IaC), CI/CD pipelines, and automated production delivery methodologies.

Have a solid background in deploying and supporting network automation or telemetry services on Kubernetes.

Experience with production on-call rotations, incident response, root-cause analysis, and driving corrective actions to completion is essential.

Essential Skills

KubernetesGoPythonGitOpsInfrastructure as CodeCI/CDProduction SupportIncident ResponseRoot Cause Analysis

Good to Have

IP RoutingData Center FabricsCloud NetworkingCluster APIMetal3Kubernetes OperatorsOpen Source Contributions

More Details

RoleSenior Platform Engineer, Network Infrastructure
DepartmentEngineering
Employment TypeFull Time, Remote

About the Company

Nvidia logo

Nvidia

IT Consulting