NCX Senior Engineer

NVIDIA

8+ yrs Bengaluru, Hyderabad, Pune Full Time Hybrid (office + remote)
NVIDIA logo
Posted : today
Actively hiring

Job description

Join NVIDIA's DSX team as an NCX Senior Engineer, focusing on the operational excellence of NVIDIA Cloud Partner (NCP) infrastructure. This role is crucial for ensuring the reliable operation of large-scale NVIDIA accelerated infrastructure in production environments.

You will guide strategic partners beyond initial deployment, focusing on advanced Day 2 operations. This includes managing infrastructure health, observability, lifecycle, rapid remediation, performance validation, and overall operational readiness. Collaboration with partner engineering and operations teams is key to developing standardized approaches for NVIDIA workloads and their broader customer ecosystems. It's a deeply technical, hands-on position at the convergence of NVIDIA accelerated computing, cloud infrastructure, distributed systems, and production operations.

Responsibilities

Lead Day 2 operational readiness for NCPs, establishing systems, procedures, and automation for consistent management of NVIDIA accelerated infrastructure.

Develop continuous infrastructure validation for GPU, CPU, storage, and network health across large AI clusters to proactively identify issues.

Enhance observability by assisting NCPs in implementing comprehensive telemetry, monitoring, alerting, and dashboards for compute, GPU, networking, storage, Kubernetes, and AI workloads.

Build automated workflows for detecting, isolating, repairing, and reintegrating unhealthy infrastructure with minimal workload disruption.

Refine fleet lifecycle administration for large GPU fleets, covering driver/firmware management, node maintenance, patching, and configuration drift.

Operationalize NVIDIA reference architectures into production practices, validation criteria, runbooks, and automation.

Define clear operational health and readiness metrics, SLOs, and validation mechanisms for infrastructure reliability.

Qualifications

Requires a BS, MS, or Ph.D. in Computer Science, Computer/Electrical Engineering, or a related technical field, or equivalent practical experience.

Possess 8+ years of experience in infrastructure engineering, SRE, DevOps, cloud platform engineering, or systems engineering supporting large-scale production environments.

Demonstrate strong experience operating Linux-based distributed systems and cloud infrastructure in production.

Exhibit a deep understanding of Kubernetes, containers, cluster scheduling, and the operational lifecycle of multi-node environments.

Have a solid grasp of production observability principles, including metrics, logging, alerting, health checks, and SLO-driven operations.

Proven experience crafting automation for infrastructure lifecycle management, failure detection, remediation, upgrades, and configuration management.

Possess strong networking fundamentals and troubleshooting skills for complex distributed systems.

Skilled in programming and automation using Python, Go, shell scripting, or similar languages.

Essential Skills

LinuxKubernetesCloud InfrastructurePythonGoShell ScriptingObservabilityInfrastructure ManagementDistributed SystemsNetworking

Good to Have

DGX/HGXCUDANVLink/NVSwitchInfiniBandRoCEGPU OperatorNetwork OperatorPrometheusGrafanaOpenTelemetryAlertmanager

Highlights

  • Actively hiring

More Details

RoleNCX Senior Engineer
IndustryTechnology, Internet, Computer Software, Computer Hardware
DepartmentOperations
Employment TypeFull Time, Hybrid (office + remote)

About the Company

Nvidia logo

Nvidia

Technology

NCX Senior Engineer at NVIDIA | SkillMX | SkillMX