Senior Staff SRE – Compute Platform
NVIDIA
NVIDIA
NVIDIA is seeking a Senior Staff SRE to architect and manage robust, scalable compute platforms essential for global engineering operations. This pivotal role involves deep engagement with Kubernetes, KubeVirt, bare-metal systems, automation strategies, observability frameworks, and AI-driven operational enhancements.
Join a dynamic team dedicated to tackling intricate infrastructure challenges, developing enduring automation solutions, and significantly elevating the reliability and operational experience of our critical compute services.
Build, maintain, and enhance large-scale compute platforms encompassing Kubernetes, KubeVirt, Linux, containers, and bare-metal infrastructure, prioritizing performance, capacity, reliability, and operational scalability.
Spearhead the end-to-end management of bare-metal provisioning in data centers, including PXE boot, DHCP, DNS, OS deployment, hardware validation, and fleet-wide automation.
Design and implement automation, self-service tools, and advanced observability solutions utilizing APIs, Python or Go, Infrastructure as Code principles, configuration management, and comprehensive metrics, logs, traces, and service health data.
Establish and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), error budgets, and alerting mechanisms. Lead incident response efforts, complex investigations, corrective actions, and foster a culture of blameless postmortems.
Collaborate closely with infrastructure, security, hardware, data-center, and application teams to drive global platform initiatives and actively participate in an on-call rotation.
A Bachelor of Science (BS) in Computer Science, Engineering, a related technical field, or equivalent practical experience, coupled with over 10 years of experience managing production infrastructure or platform services.
Demonstrate strong expertise in Kubernetes administration, KubeVirt, Docker, containerization technologies, microservices architecture, Linux systems, and proficiently resolving complex distributed-system challenges.
Possess hands-on experience deploying and managing bare-metal infrastructure within data center environments, covering provisioning, networking, operating-system lifecycle management, and hardware automation.
Exhibit proficiency in programming languages such as Python or Go, or a comparable language, with a proven track record of building RESTful services and integrating infrastructure APIs.
Showcase experience with Infrastructure as Code and automation tools like Terraform, Ansible, Chef, or Puppet, alongside a firm grasp of TCP/IP networking and infrastructure security best practices.
Bring substantial Site Reliability Engineering (SRE) and observability expertise, including a deep understanding of SLIs, SLOs, error budgets, incident management, monitoring, logging, tracing, and familiarity with tools like OpenTelemetry, Prometheus, Grafana, ELK Stack, or Splunk.
Possess exceptional written and interpersonal communication skills, substantiated by a history of delivering practical, scalable solutions to complex technical problems.
Nvidia
IT Consulting