Senior Staff SRE – Compute Platform

NVIDIA

10+ yrs Bengaluru Full Time Hybrid (office + remote)
NVIDIA logo
Posted : today
Actively hiring

Job description

NVIDIA is seeking a Senior Staff SRE to architect and manage robust, scalable compute platforms essential for global engineering operations. This pivotal role involves deep engagement with Kubernetes, KubeVirt, bare-metal systems, automation strategies, observability frameworks, and AI-driven operational enhancements.

Join a dynamic team dedicated to tackling intricate infrastructure challenges, developing enduring automation solutions, and significantly elevating the reliability and operational experience of our critical compute services.

Responsibilities

Build, maintain, and enhance large-scale compute platforms encompassing Kubernetes, KubeVirt, Linux, containers, and bare-metal infrastructure, prioritizing performance, capacity, reliability, and operational scalability.

Spearhead the end-to-end management of bare-metal provisioning in data centers, including PXE boot, DHCP, DNS, OS deployment, hardware validation, and fleet-wide automation.

Design and implement automation, self-service tools, and advanced observability solutions utilizing APIs, Python or Go, Infrastructure as Code principles, configuration management, and comprehensive metrics, logs, traces, and service health data.

Establish and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), error budgets, and alerting mechanisms. Lead incident response efforts, complex investigations, corrective actions, and foster a culture of blameless postmortems.

Collaborate closely with infrastructure, security, hardware, data-center, and application teams to drive global platform initiatives and actively participate in an on-call rotation.

Qualifications

A Bachelor of Science (BS) in Computer Science, Engineering, a related technical field, or equivalent practical experience, coupled with over 10 years of experience managing production infrastructure or platform services.

Demonstrate strong expertise in Kubernetes administration, KubeVirt, Docker, containerization technologies, microservices architecture, Linux systems, and proficiently resolving complex distributed-system challenges.

Possess hands-on experience deploying and managing bare-metal infrastructure within data center environments, covering provisioning, networking, operating-system lifecycle management, and hardware automation.

Exhibit proficiency in programming languages such as Python or Go, or a comparable language, with a proven track record of building RESTful services and integrating infrastructure APIs.

Showcase experience with Infrastructure as Code and automation tools like Terraform, Ansible, Chef, or Puppet, alongside a firm grasp of TCP/IP networking and infrastructure security best practices.

Bring substantial Site Reliability Engineering (SRE) and observability expertise, including a deep understanding of SLIs, SLOs, error budgets, incident management, monitoring, logging, tracing, and familiarity with tools like OpenTelemetry, Prometheus, Grafana, ELK Stack, or Splunk.

Possess exceptional written and interpersonal communication skills, substantiated by a history of delivering practical, scalable solutions to complex technical problems.

Essential Skills

KubernetesKubeVirtLinuxContainerizationMicroservicesPythonGoInfrastructure as CodeTerraformAnsibleChefPuppetTCP/IP NetworkingInfrastructure SecuritySite Reliability Engineering (SRE)ObservabilitySLIsSLOsIncident ManagementMonitoringLoggingTracingOpenTelemetryPrometheusGrafanaELK StackSplunkRESTful APIs

Good to Have

HPCAIGPU-accelerated computingBare-metal compute infrastructureGPU-enabled KubernetesGPU-enabled KubeVirtVMware vSphereRed Hat OpenShiftKVMFirecrackerOpenStackNutanix AHVGenerative AIAgentic workflowsRBACService accountsSecrets managementAudit controlsWorkflow orchestrationIncident-management systems

Highlights

  • Actively hiring

More Details

RoleSenior Staff SRE – Compute Platform
DepartmentSite Reliability Engineering
Employment TypeFull Time, Hybrid (office + remote)

About the Company

Nvidia logo

Nvidia

IT Consulting

Senior Staff SRE – Compute Platform at NVIDIA | SkillMX | SkillMX