Senior Staff Site Reliability Engineer
NVIDIA
NVIDIA
Join NVIDIA, a pioneer in computer graphics, PC gaming, and accelerated computing for three decades. We are driving the next era of computing with AI, where our GPUs power intelligent systems like generative AI and self-driving cars. As a NVIDIAN, you'll thrive in a diverse and supportive environment, contributing to groundbreaking work.
We are seeking a seasoned Senior Staff Site Reliability Engineer to enhance our infrastructure's efficiency and performance across on-premise and cloud environments. This role is crucial in our mission to lead technological innovation.
You will lead the transformation of our IT Compute Core Team architecture to develop new service offerings for both on-premise and cloud platforms. This involves designing, scaling, and deploying essential infrastructure services such as DNS, NTP/PTP, DHCP, and LDAP, ensuring global-scale performance and reliability.
Key responsibilities include defining metrics for service efficiency, driving optimizations through software and hardware, and leveraging technologies like eBPF and XDP for enhanced observability and DDoS mitigation. You will also analyze system data for capacity planning, develop enterprise-wide system plans, and create tools for data collection, analysis, and visualization.
Collaboration with NVIDIA leadership, senior engineers, and product managers is essential to deliver compelling IT products and services tailored to customer needs. This includes experience with containerization and distributed systems infrastructure.
A Bachelor's degree in Engineering, Computer Science, Mathematics, or a related field, or equivalent practical experience, is required. We are looking for at least 15 years of experience in compute platform engineering, with a strong emphasis on automation and a proven track record in designing and deploying containerization architectures and distributed systems infrastructure.
Essential skills include strong analytical abilities, proficiency in developing tools for data analysis and performance profiling using tools like Terraform and configuration management, and expertise in programming languages such as Go and/or Python. Deep understanding of Linux OS, kernel internals, and experience managing large-scale bare-metal build infrastructures are critical.
Knowledge of network protocols and architectures (VLAN/VxLAN/SDN/BGP/Anycast) is also a must. Standing out will be advantageous with deep insights into DNS, LDAP, and security tools, alongside hands-on experience with container deployment and management, microservices architecture, and infrastructure as code.
Nvidia
IT Consulting