Join Google Cloud's elite team of software engineers, driving the next generation of technologies that transform how billions connect and interact with information. We seek innovative minds eager to contribute to large-scale system design, distributed computing, AI, and more. This role offers the chance to work on projects vital to Google Cloud's success, with opportunities to evolve and adapt within a fast-paced environment. You'll anticipate user needs, act with ownership, and champion innovation, demonstrating leadership and enthusiasm for tackling complex, full-stack challenges.
Our team is dedicated to enhancing the machine learning software and hardware stack, particularly focusing on the TPU compiler for both internal Google teams and external Google Cloud Platform (GCP) customers. This is an opportunity to build the infrastructure that powers advanced machine learning and shape the future of TPU performance.
Key responsibilities include gaining a deep understanding of Google's ML stack, including frameworks like JAX and PyTorch, along with the XLA and runtime stack.
You will research and develop innovative compiler optimizations tailored for ML workloads and emerging architectures. Identifying opportunities to boost ML workload efficiency through insightful performance debugging and custom kernel development will be crucial, leading to the creation of effective compiler solutions.
Furthermore, this role involves providing technical leadership and mentorship, potentially as a Team Lead (TL), and exploring strategic initiatives to advance our capabilities.
A Bachelor's degree or equivalent practical experience is required, alongside a minimum of 8 years in software development and 3 years in software design and architecture.
Proficiency in machine learning, compilers, computer architecture, GPU programming, C++, and Python is essential.
Preferred qualifications include experience with open-source software development, ML compiler internals, writing optimization passes, and debugging correctness and performance issues across the ML software stack. Familiarity with accelerator hardware architectures like TPUs and GPUs, and experience with performance analysis on these systems, will be highly beneficial.
IT Consulting