Senior Staff Site Reliability Engineer
NVIDIA
NVIDIA
NVIDIA, a pioneer in transforming computer graphics, PC gaming, and accelerated computing, is leveraging AI to usher in a new era of computing. Join our team in India as a Senior Staff Site Reliability Engineer and play a pivotal role in our AI-powered enterprise platforms. This role is critical for shaping the future of reliability, combining incident leadership with deep engineering expertise.
As an NVIDIAN, you'll be part of a diverse and supportive environment where innovation thrives. We are committed to pushing the boundaries of what's possible, and your contributions will have a lasting global impact. Embrace the opportunity to define the next era of computing alongside the world's best talent.
Lead critical incidents from start to finish, including triage, cross-team coordination, and executive communication across global time zones. Define and drive SRE initiatives that enhance reliability, scalability, and developer efficiency for NVIDIA's enterprise systems.
Design, build, and manage distributed systems, including Kubernetes and cloud-native infrastructure, powering our AI enterprise products. Develop automation for incident detection, communication, and remediation, transitioning manual processes to self-healing systems.
Enhance observability to improve early detection and reduce alert noise. Conduct thorough root cause analyses, translating learnings into systemic fixes and preventative measures. Apply AI and data-driven techniques to optimize incident triage, summarization, and decision support. Champion AI-assisted engineering practices to streamline daily workflows. Partner with Cloud, Platform, Security, and AI/ML teams to embed SRE best practices, define SLOs, and influence architectural decisions early.
Mentor engineers, elevate engineering standards through reviews, and foster a robust reliability culture within our India organization.
We are seeking a seasoned professional with over 10 years of experience in Site Reliability Engineering, Production Engineering, Platform Engineering, or Incident Management, demonstrating strong technical leadership. A BS or MS degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience, is required.
Proven experience as an Incident Commander or leading major incident responses in complex, high-availability settings is essential. A deep understanding of distributed systems, monitoring, and reliability engineering principles, including SLIs/SLOs, error budgets, capacity planning, and graceful degradation, is critical.
Proficiency in at least one programming language (e.g., Python, Go, Java) for building production-grade automation is necessary, alongside hands-on expertise with public cloud platforms (AWS, Azure, GCP) and container technologies like Docker and Kubernetes. Solid experience with infrastructure-as-code tools (e.g., Terraform, AWS CDK, CloudFormation) and CI/CD pipelines is expected.
Strong fundamentals in Linux/Unix and networking, coupled with expertise in observability tools such as OpenTelemetry, Prometheus, and Grafana, are key. Working knowledge of relational databases (e.g., PostgreSQL, MySQL), including SQL, indexing, and query optimization, is beneficial. Familiarity with AI/ML concepts applied to operational workflows is a plus.
Exceptional written and verbal communication skills are vital for briefing executives during high-pressure incidents and influencing senior technical stakeholders.
Nvidia
Information Technology & Services