Join Google Cloud's dynamic team and contribute to the evolution of next-generation technologies that connect billions globally. We are seeking versatile engineers with innovative ideas from diverse fields, including AI, distributed systems, and large-scale design. As a software engineer, you will engage in impactful projects critical to Google Cloud's success, with opportunities for growth and team mobility. Empowered to act as an owner, you will drive innovation and anticipate customer needs.
Our team is central to Google's Core ML organization, providing essential ML software tools and hardware infrastructure across all Google product areas. We are at the forefront of advancing ML excellence for Google and the world. From ML frameworks and performance debugging to ML efficiency and applied ML, our work spans critical aspects of AI and infrastructure, pushing the boundaries of ML capabilities.
Google Cloud empowers organizations with enterprise-grade solutions, leveraging Google's advanced technology to accelerate digital transformation. We provide tools for sustainable development, enabling businesses in over 200 countries to achieve growth and overcome complex challenges.
Drive continuous improvements across Google's ML software and hardware stack, benefiting both internal teams and Google Cloud Platform customers.
Develop a comprehensive understanding of Google's ML training and serving infrastructure, with a focus on Frameworks like JAX and PyTorch, Accelerated Linear Algebra (XLA), and the runtime stack.
Identify and implement solutions for enhancing ML workload efficiency through in-depth performance debugging of custom kernels and workloads.
Collaborate with cross-functional teams owning different ML stack components to address performance optimization use cases. Provide insights into performance bottlenecks within open-source ML inference frameworks such as TorchTPU, vLLM, and SGLang.
Support emerging ML paradigms, including horizontal scaling for advanced TPU chips, by contributing to the ML stack and performance analysis tools.
A Bachelor's degree or equivalent practical experience is required.
Possess at least 5 years of professional software development experience, proficient in C++ or Python.
Demonstrate experience in machine learning (ML) infrastructure development or ML performance engineering.
Preferred qualifications include expertise in ML compilers and their internal workings, experience with writing compiler optimization passes, and familiarity with accelerator hardware architectures like TPUs and GPUs. Experience with ML inference frameworks (e.g., vLLM, SG Lang, Pathways) and ML frameworks (e.g., TensorFlow, JAX, PyTorch, Keras) is also highly valued.
Cloud Computing