Server Performance Architect - Hardware
NVIDIA
NVIDIA
Join NVIDIA, a leader in technological innovation, and shape the future of AI server systems. We're looking for talented architects to drive next-generation server performance, combining deep architectural insight with hands-on silicon expertise. This role involves analyzing complex workloads, investigating system performance on both NVIDIA and competitor platforms, and developing specialized benchmarks to pinpoint architectural strengths and weaknesses.
At NVIDIA, you'll be part of a dynamic, learning-oriented environment where challenging problems are solved. Your work will directly impact the scalability and efficiency of AI computation, amplifying human creativity and intelligence. We foster a supportive and diverse team culture, encouraging everyone to achieve their best.
Define and advance server-level performance objectives across critical subsystems including CPU, GPU, memory, interconnects, networking, and storage. Conduct in-depth workload characterization and bottleneck analysis on diverse server platforms using industry-standard AI, inference, and HPC benchmarks. Employ advanced profiling and tracing tools to diagnose performance issues and identify system-level optimization opportunities.
Perform detailed trade-off studies on system topology, thermal and power envelopes, and memory hierarchies to inform key architectural decisions. Collaborate closely with silicon, platform, firmware, and software engineering teams to resolve performance gaps from initial bring-up through production. Develop robust automation and tooling for tracking and reporting performance regressions. Represent the performance perspective effectively in architecture reviews and cross-functional design discussions. Build and maintain analytical models and simulation frameworks for future server platforms.
A Bachelor's or Master's degree in Electrical/Computer Engineering, Computer Science, or a related field is required, alongside a minimum of 10 years in server/system performance architecture. A strong grasp of modern server architectures, including CPU microarchitecture, PCIe/CXL, memory subsystems (DDR/HBM), and coherency protocols, is essential. Proven hands-on experience with system-level profiling and performance analysis tools on server platforms is necessary.
Solid knowledge of GPU-accelerated computing, high-performance networking, or high-performance storage subsystems is expected. Proficiency in scripting languages like Python or C/C++ for tool development and data analysis is required. You should be comfortable leveraging AI-powered tools to enhance productivity and workflow efficiency. Excellent communication skills are crucial for translating complex performance data into clear, actionable architectural recommendations.
Nvidia
Technology