Senior Staff Site Reliability Engineer
NVIDIA
NVIDIA
NVIDIA is seeking a Senior Staff Software Engineer to architect and lead the development of the runtime foundation for its enterprise AI platforms. This role involves providing technical leadership for systems that deploy, operate, and scale AI applications, inference services, and databases across diverse cloud and on-premises environments. Join a team dedicated to pushing the boundaries of AI and shaping the future of computing.
As a leading technology company with a rich history of innovation, NVIDIA fosters a supportive and diverse environment where employees are inspired to achieve their best. We are committed to equal opportunity and building a workplace that reflects the diversity of the world.
Define the architectural roadmap for scalable enterprise AI runtime platforms. Design and implement Kubernetes-based systems for deploying and scaling AI applications and services. Develop control-plane services, APIs, and automation for workload management, including provisioning, configuration, and recovery. Enhance runtime capabilities for GPU scheduling, autoscaling, load balancing, and rate limiting. Improve the performance, availability, and developer experience of large-scale AI inference services. Build and automate robust database services, including relational and vector databases, ensuring high availability and failover. Establish secure and consistent application lifecycle management patterns for both cloud and on-premises deployments. Develop advanced observability tools for comprehensive monitoring, profiling, and debugging of applications and resources. Lead cross-functional technical initiatives, mentor engineers, and set standards for long-term platform evolution.
A Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience is required. Candidates should possess 8+ years of experience in software engineering, with a strong background in distributed systems, cloud infrastructure, database platforms, or large-scale backend services.
Essential skills include proficiency in Python, Go, C++, or Java, with a proven ability to deliver production-grade systems. Demonstrated experience in designing scalable, highly available Kubernetes-based platforms is crucial. The ideal candidate will have a track record of leading technical strategy, influencing cross-team efforts, and solving complex platform challenges.
Experience in building control planes, platform APIs, Kubernetes operators, or workload lifecycle management systems is necessary. Familiarity with high-performance services for AI inference or low-latency workloads is expected. Practical knowledge of relational or vector databases, including aspects like availability, replication, and performance tuning, is important. Experience with GitOps, CI/CD, observability, and cloud-native security practices is also required.
Nvidia
Technology