Senior Network Site Reliability Engineer
NVIDIA
NVIDIA
Join our Enterprise Network Operations and SRE team as a Senior Network Site Reliability Engineer. You will play a vital role in shaping our vision for a robust and efficient network infrastructure. This position is ideal for individuals passionate about network operations and dedicated to elevating the user experience.
This role offers the chance to tackle complex network challenges through hands-on debugging, with a strong emphasis on network automation, observability, comprehensive documentation, and operational excellence. Your contributions will be crucial in ensuring exceptional user satisfaction and driving brilliance in our network operations.
You will own the operational integrity of the network infrastructure, guaranteeing high availability and reliability. This includes actively managing network incidents and service requests.
Collaborate with architecture and deployment teams to ensure new implementations are supportable and adhere to production standards.
Champion and implement automation initiatives to minimize manual tasks, thereby enhancing operational efficiency and achieving Service Level Objectives (SLOs).
Monitor network performance, identify areas for enhancement, and work with relevant teams to implement improvements. Proactively mitigate network risks for continuous advancement.
Partner with cross-functional domain experts to swiftly resolve production issues and ensure customer satisfaction. Conduct blameless postmortems and follow through on Root Cause Analyses (RCAs).
Identify opportunities for operational improvements and collaborate with colleagues to develop solutions that boost excellence and sustainability in network operations. Develop knowledge base articles for automation and bots.
A Bachelor's degree in Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience, is required.
Possess a minimum of 10 years of industry experience in network operations or related fields, with a focus on support, automation, and site reliability engineering. Familiarity with both enterprise and data center networks is essential.
Demonstrate strong proficiency in network fundamentals and experience resolving complex network issues, including expertise in technologies like TCP/UDP, IPv4/IPv6, BGP, OSPF, ISIS, VPN, L2 switching, Firewalls, Load Balancers, Data Center Network technologies, and Wireless. A consistent track record in network operations is expected.
Familiarity with network management tools such as Prometheus, Grafana, Alert Manager, Nautobot/Netbox, and BigPanda is necessary.
Experience automating networks using frameworks like Salt, Ansible, Python, or similar is required.
Skills with ServiceNow, Jira, and foundational knowledge of the ITIL framework are beneficial.
Knowledge of Linux system fundamentals is important.
Possess a detailed problem-solving approach, critical thinking abilities, strong interpersonal skills, and a solid understanding of ownership and drive.
Nvidia
IT Consulting