Lead the evolution of AI/ML infrastructure and performance engineering at Google. This role involves technical leadership and people management, guiding a team of engineers to optimize both their code and the systems they build. Contribute to product strategy and team development, working on cutting-edge projects in areas like artificial intelligence, natural language processing, and large-scale system design.
As a manager, you will oversee project goals, manage a substantial product budget, and direct the deployment of large-scale international projects. This position requires a strong blend of technical acumen and leadership to drive innovation and ensure the efficiency and reliability of Google's AI/ML platforms.
Gain a deep understanding of current data collection, analysis, and visualization workflows, with a focus on frameworks, XLA, and runtime stacks.
Drive advancements in ML paradigms, such as horizontal scaling for new TPU chips, by contributing to the entire stack and analysis tools.
Collaborate with ML Stack leads to identify model optimization use cases and integrate debugging capabilities seamlessly into third-party environments like VSCode, Cursor, and Grafana.
Analyze and provide insights into performance improvements for OSS ML inference frameworks such as vLLM and SGLang using Xprof.
Partner with teams responsible for various ML stack components to identify performance optimization opportunities, working with frameworks like TorchTPU, vLLM, and SGLang to pinpoint bottlenecks.
A Bachelor's degree or equivalent practical experience is required.
Possess at least 8 years of experience spanning software engineering, machine learning infrastructure, computer architecture, distributed computing, and people management.
Demonstrated experience in people management and building infrastructure to enhance the performance of Machine Learning (ML) systems and applications is essential.
Preferred qualifications include experience with Performance Optimization, GPU Programming, High Performance Computing, Large Language Models, and open-source contributions. Hands-on experience with ML frameworks like TensorFlow, JAX, PyTorch, and Keras, alongside ML Inference frameworks such as vLLM and SG Lang, is highly advantageous. Experience in building agentic workflows for performance debugging and optimization is also a plus.
IT Consulting