Senior Network Site Reliability Engineer

NVIDIA

10+ yrs Bengaluru, Hyderabad, Pune Full Time Hybrid (office + remote)
NVIDIA logo
Posted : today
Actively hiring

Job description

Join our Enterprise Network Operations and SRE team as a Senior Network Site Reliability Engineer, instrumental in shaping a resilient and efficient network infrastructure. This role is perfect for individuals passionate about network operations and dedicated to elevating user experience. You will tackle complex network challenges through hands-on debugging, with a strong emphasis on network automation, observability, documentation, and operational excellence. Your contributions will be key to ensuring user satisfaction and driving brilliance in network operations.

Responsibilities

You will own the operational integrity of the network infrastructure, ensuring high availability and reliability, while actively managing network incidents and service requests. Collaborate with architecture and deployment teams to validate that new implementations are supportable and adhere to production standards.

Drive operational efficiency by advocating for and implementing automation, minimizing manual tasks to meet and maintain Service Level Objectives (SLOs). Monitor network performance, identify improvement areas, and implement refinements in coordination with relevant teams. Proactively mitigate network risks to foster continuous improvement.

Work with cross-functional domain experts to resolve production issues swiftly and effectively, ensuring customer happiness. Conduct blameless postmortems and follow through on Root Cause Analyses (RCAs). Identify operational improvement opportunities and collaborate with colleagues on solutions that enhance excellence and sustainability in network operations. Develop knowledge base articles for automation and bots.

Qualifications

A Bachelor's degree in Computer Science, Electrical Engineering, or a related technical field, or equivalent experience, is required. We are looking for a minimum of 10 years of industry experience in network operations or related fields, focusing on support, automation, and site reliability engineering. Familiarity with both enterprise and data center networks is crucial.

Possess strong proficiency in network fundamentals and a proven ability to fix complex network issues, with expertise in technologies such as TCP/UDP, IPv4/IPv6, BGP, OSPF, ISIS, VPN, L2 switching, Firewalls, Load Balancers, Data Center Network technologies, and Wireless. A consistent track record in network operations is essential.

Familiarity with network management tools like Prometheus, Grafana, Alert Manager, Nautobot/Netbox, and BigPanda is necessary. Experience in network automation using frameworks such as Salt, Ansible, or Python is required, along with skills in ServiceNow, Jira, and foundational knowledge of the ITIL framework. Understanding of Linux system fundamentals is also important. Candidates should demonstrate a detailed problem-solving approach, critical thinking, strong interpersonal skills, and a solid sense of ownership and drive.

Essential Skills

Network OperationsSite Reliability EngineeringTCP/UDPIPv4IPv6BGPOSPFISISVPNL2 switchingFirewallsLoad BalancersData Center Network technologiesWirelessPrometheusGrafanaAlert ManagerNautobot/NetboxBigPandaSaltAnsiblePythonServiceNowJiraLinux System FundamentalsProblem-SolvingCritical ThinkingInterpersonal SkillsOwnershipDrive

Good to Have

SNMPSyslogStreaming TelemetryMellanox/Cumulus LinuxCisco/AristaPalo Alto firewallsVersa SDWANNetscalersF5 load balancersGoVXLAN/EVPNMPLSRSVPSegment RoutingSDWANSASE Platforms

Highlights

  • Actively hiring

More Details

RoleSenior Network Site Reliability Engineer
DepartmentSite Reliability Engineering
Employment TypeFull Time, Hybrid (office + remote)

About the Company

Nvidia logo

Nvidia

IT Consulting

Senior Network Site Reliability Engineer at NVIDIA | SkillMX | SkillMX