Senior Staff Site Reliability Engineer

NVIDIA

10+ yrs Bengaluru, Pune Full Time Hybrid (office + remote)
NVIDIA logo
Posted : today
Actively hiring

Job description

NVIDIA, a pioneer in transforming computer graphics, PC gaming, and accelerated computing, is leveraging AI to usher in a new era of computing. Join our team in India as a Senior Staff Site Reliability Engineer and play a pivotal role in our AI-powered enterprise platforms. This role is critical for shaping the future of reliability, combining incident leadership with deep engineering expertise.

As an NVIDIAN, you'll be part of a diverse and supportive environment where innovation thrives. We are committed to pushing the boundaries of what's possible, and your contributions will have a lasting global impact. Embrace the opportunity to define the next era of computing alongside the world's best talent.

Responsibilities

Lead critical incidents from start to finish, including triage, cross-team coordination, and executive communication across global time zones. Define and drive SRE initiatives that enhance reliability, scalability, and developer efficiency for NVIDIA's enterprise systems.

Design, build, and manage distributed systems, including Kubernetes and cloud-native infrastructure, powering our AI enterprise products. Develop automation for incident detection, communication, and remediation, transitioning manual processes to self-healing systems.

Enhance observability to improve early detection and reduce alert noise. Conduct thorough root cause analyses, translating learnings into systemic fixes and preventative measures. Apply AI and data-driven techniques to optimize incident triage, summarization, and decision support. Champion AI-assisted engineering practices to streamline daily workflows. Partner with Cloud, Platform, Security, and AI/ML teams to embed SRE best practices, define SLOs, and influence architectural decisions early.

Mentor engineers, elevate engineering standards through reviews, and foster a robust reliability culture within our India organization.

Qualifications

We are seeking a seasoned professional with over 10 years of experience in Site Reliability Engineering, Production Engineering, Platform Engineering, or Incident Management, demonstrating strong technical leadership. A BS or MS degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience, is required.

Proven experience as an Incident Commander or leading major incident responses in complex, high-availability settings is essential. A deep understanding of distributed systems, monitoring, and reliability engineering principles, including SLIs/SLOs, error budgets, capacity planning, and graceful degradation, is critical.

Proficiency in at least one programming language (e.g., Python, Go, Java) for building production-grade automation is necessary, alongside hands-on expertise with public cloud platforms (AWS, Azure, GCP) and container technologies like Docker and Kubernetes. Solid experience with infrastructure-as-code tools (e.g., Terraform, AWS CDK, CloudFormation) and CI/CD pipelines is expected.

Strong fundamentals in Linux/Unix and networking, coupled with expertise in observability tools such as OpenTelemetry, Prometheus, and Grafana, are key. Working knowledge of relational databases (e.g., PostgreSQL, MySQL), including SQL, indexing, and query optimization, is beneficial. Familiarity with AI/ML concepts applied to operational workflows is a plus.

Exceptional written and verbal communication skills are vital for briefing executives during high-pressure incidents and influencing senior technical stakeholders.

Essential Skills

Site Reliability EngineeringProduction EngineeringPlatform EngineeringIncident ManagementTechnical LeadershipDistributed SystemsKubernetesCloud-Native InfrastructureAutomationObservabilityPythonGoJavaAWSAzureGCPDockerTerraformCI/CDLinuxUnixNetworkingOpenTelemetryPrometheusGrafanaSQLPostgreSQLMySQLLLMsAnomaly DetectionCommunication

Good to Have

AI/MLChatOpsWorkflow OrchestrationAlert IntelligenceMesaOpen Source ContributionsConference Talks

Highlights

  • Actively hiring

More Details

RoleSenior Staff Site Reliability Engineer
IndustryInformation Technology & Services, Semiconductors
DepartmentOperations
Employment TypeFull Time, Hybrid (office + remote)

About the Company

Nvidia logo

Nvidia

Information Technology & Services

Senior Staff Site Reliability Engineer at NVIDIA | SkillMX | SkillMX