Senior Site Reliability Engineering, Storage
NVIDIA
NVIDIA
Join NVIDIA and be at the forefront of AI-driven computing. We are seeking a Senior Site Reliability Engineer specializing in Storage to ensure the utmost reliability, performance, and scalability of our global storage platforms. These systems are critical for both internal and external services, requiring a blend of deep storage expertise and robust SRE practices.
This role offers an opportunity to shape the future of computing infrastructure at NVIDIA, an industry leader with a legacy of innovation. You'll work in a supportive and diverse environment, contributing to defining the next era of technology.
Lead the design, deployment, and operation of production NAS, SAN, and Object Storage platforms, ensuring peak reliability, performance, and security.
Gather requirements from partner teams, architect sophisticated storage solutions, and oversee their end-to-end implementation for existing and new services.
Develop and maintain advanced automation for provisioning, configuration, monitoring, incident response, and lifecycle management of storage infrastructure.
Participate in on-call rotations, lead complex troubleshooting efforts for storage and performance issues, and drive thorough root cause analysis and preventative measures.
Define, monitor, and analyze SLOs/SLIs and error budgets for storage services, leveraging observability and analytics for continuous improvement.
Create and maintain essential runbooks, standard operating procedures, and comprehensive documentation for storage services and automation.
Analyze capacity and usage trends, forecast future needs, and recommend strategies for scaling and optimization to support business objectives.
Collaborate effectively with SRE, infrastructure, networking, and application teams in a follow-the-sun model to deliver consistent, high-quality service globally.
Mentor junior engineers, promote best practices, and champion the adoption of SRE principles across the team.
A minimum of 12 years of experience in Site Reliability, DevOps, or Infrastructure Engineering, with a significant emphasis on storage systems.
A Bachelor’s degree in Computer Science, Computer Engineering, or a related technical field, or equivalent practical experience.
Demonstrated hands-on experience in designing, deploying, and operating enterprise-grade NAS, SAN, and/or Object Storage platforms.
A strong understanding of core SRE concepts, including SLOs/SLIs, error budgets, incident management, observability, and postmortems.
Proficiency with Infrastructure as Code and configuration management tools such as Terraform, Ansible, Puppet, and SaltStack, along with experience using source control systems.
Experience building and operating highly available, scalable infrastructure, including automation for provisioning, monitoring, and remediation.
Familiarity with container and virtualization platforms like Docker, Kubernetes, and hypervisors, as well as modern CI/CD and version control tools.
Solid scripting or programming skills in languages like Python, Go, or Shell for tool development, workflow automation, and system integration.
Excellent communication and collaboration skills are essential for working effectively across distributed and cross-functional teams.
Nvidia
Semiconductors