Join Google Cloud as a Staff ML Compiler Engineer, focusing on TPU Performance Optimizations. This role is crucial for advancing the performance of machine learning software and hardware stacks, specifically for Google's internal teams and external Google Cloud Platform (GCP) customers.
You will play a key part in developing the infrastructure that powers advanced machine learning, directly shaping the future of Tensor Processing Unit (TPU) performance. Your work will involve designing and implementing sophisticated compiler optimizations like pipelining and fusions to maximize TPU efficiency.
This position offers the opportunity to influence significant Alphabet initiatives, including Large Language Model (LLM) development and chip co-design. Additionally, you'll empower external cloud customers and collaborate across teams to integrate frameworks such as PyTorch, ensuring our hardware remains a leading Machine Learning (ML) platform.
Engage with and develop a deep understanding of various components within Google's ML stack, including frameworks like JAX and PyTorch, as well as the XLA and runtime stack.
Conduct research and develop innovative compiler optimizations tailored for ML workloads and novel architectures.
Identify and implement enhancements for ML workload efficiency by performing insightful performance debugging on ML workloads and custom kernels, and developing corresponding compiler solutions.
Provide leadership and mentorship, potentially in a Team Lead (TL) capacity, and contribute to strategic initiatives.
A Bachelor's degree or equivalent practical experience is required, coupled with 8 years of software development experience and 3 years in software design and architecture.
Proficiency is expected in machine learning, compilers, computer architecture, GPU programming, C++, and Python.
Preferred qualifications include experience with open-source software development, releasing and supporting open-source projects, and hands-on experience with ML compilers and their internals, including writing optimization passes.
Demonstrated experience in debugging correctness and performance issues across the entire ML Software (SW) stack is highly valued. Familiarity with accelerator hardware architectures (TPUs/GPUs) and performance analysis for these systems is also beneficial.
IT Consulting