Architect - GPU Performance
NVIDIA
NVIDIA
Join NVIDIA, a leader in GPU innovation, and contribute to the advancement of next-generation visual computing, automotive, GPU, and HPC systems.
As an Architect on our team, you will immerse yourself in a supportive and diverse environment, dedicated to tackling complex challenges. You will play a crucial role in shaping the future of high-performance computing by working on cutting-edge CPU and memory subsystems, next-gen GPUs, and interconnect fabrics.
This is an opportunity to amplify human creativity and intelligence through groundbreaking technology. Embrace a role where your contributions make a lasting global impact.
Perform system-level performance and bottleneck analysis for complex, high-performance GPUs and System-on-Chips (SoCs).
Engage with hardware models across various abstraction levels, including performance models, RTL testbenches, emulators, and silicon, to identify performance bottlenecks.
Understand and define key product performance use-cases. Develop tailored workloads and test suites for graphics, machine learning, automotive, video, and compute vision applications.
Collaborate closely with architecture and design teams to evaluate architectural trade-offs impacting system performance, area, and power consumption.
Develop essential infrastructure such as performance models, testbench components, and performance analysis/visualization tools.
A Bachelor's or Master's degree in a relevant field is required, with a PhD being a plus, or equivalent experience.
Possess at least 3 years of experience focused on performance analysis within complex System-on-Chip (SoC) and/or GPU architectures.
Demonstrate a strong command of SoC architecture, graphics pipelines, memory subsystem architecture, and Network-on-Chip (NoC)/Interconnect architecture.
Exhibit expert proficiency in programming (C/C++) and scripting languages (Perl/Python). Familiarity with Verilog/System Verilog, SystemC/TLM is highly advantageous.
Possess robust debugging and analysis skills, including data and statistical analysis, and the ability to debug failures using RTL dumps.
Hands-on experience developing performance simulators or cycle-accurate/approximate models for pre-silicon performance analysis is a significant advantage.
Nvidia
Semiconductors