Senior Network Site Reliability Engineer
NVIDIA
NVIDIA
Join our Enterprise Network Operations and SRE team as a Senior Network Site Reliability Engineer, instrumental in shaping a resilient and efficient network infrastructure. This role is perfect for individuals passionate about network operations and dedicated to elevating user experience. You will tackle complex network challenges through hands-on debugging, with a strong emphasis on network automation, observability, documentation, and operational excellence. Your contributions will be key to ensuring user satisfaction and driving brilliance in network operations.
You will own the operational integrity of the network infrastructure, ensuring high availability and reliability, while actively managing network incidents and service requests. Collaborate with architecture and deployment teams to validate that new implementations are supportable and adhere to production standards.
Drive operational efficiency by advocating for and implementing automation, minimizing manual tasks to meet and maintain Service Level Objectives (SLOs). Monitor network performance, identify improvement areas, and implement refinements in coordination with relevant teams. Proactively mitigate network risks to foster continuous improvement.
Work with cross-functional domain experts to resolve production issues swiftly and effectively, ensuring customer happiness. Conduct blameless postmortems and follow through on Root Cause Analyses (RCAs). Identify operational improvement opportunities and collaborate with colleagues on solutions that enhance excellence and sustainability in network operations. Develop knowledge base articles for automation and bots.
A Bachelor's degree in Computer Science, Electrical Engineering, or a related technical field, or equivalent experience, is required. We are looking for a minimum of 10 years of industry experience in network operations or related fields, focusing on support, automation, and site reliability engineering. Familiarity with both enterprise and data center networks is crucial.
Possess strong proficiency in network fundamentals and a proven ability to fix complex network issues, with expertise in technologies such as TCP/UDP, IPv4/IPv6, BGP, OSPF, ISIS, VPN, L2 switching, Firewalls, Load Balancers, Data Center Network technologies, and Wireless. A consistent track record in network operations is essential.
Familiarity with network management tools like Prometheus, Grafana, Alert Manager, Nautobot/Netbox, and BigPanda is necessary. Experience in network automation using frameworks such as Salt, Ansible, or Python is required, along with skills in ServiceNow, Jira, and foundational knowledge of the ITIL framework. Understanding of Linux system fundamentals is also important. Candidates should demonstrate a detailed problem-solving approach, critical thinking, strong interpersonal skills, and a solid sense of ownership and drive.
Nvidia
IT Consulting