Software Engineering Manager, AI/ML Infrastructure and Performance Engineering

Google

8 yrs Bengaluru Full Time Hybrid (office + remote)
Google logo
Posted : today
Actively hiring

Job description

Lead and inspire a team of engineers at the forefront of AI/ML infrastructure and performance engineering. This role offers the opportunity to shape technical strategy, drive major projects, and foster the growth of your team. You will be instrumental in optimizing not only your own code but also empowering other engineers to achieve peak performance.

Our work spans critical areas like information retrieval, artificial intelligence, natural language processing, and large-scale system design. We operate at immense scale and speed, and as a manager, you will guide the direction of our exceptional software engineers, ensuring they are empowered to innovate and excel.

This position involves managing engineers across diverse teams and locations, overseeing substantial product budgets, and directing the international deployment of large-scale projects. It’s a chance to make a significant impact on Google's core services and cloud offerings.

Responsibilities

Drive continuous improvements in machine learning software and hardware stacks by providing in-depth performance debugging for workloads and custom kernels. Summarize insights from various profile data views, including trace timelines, memory usage, compiler profiles, and ML graph summaries.

Contribute across the end-to-end stack and analysis tools to support new ML paradigms, such as horizontal scaling for upcoming TPU chips. Collaborate with ML Stack leads to grasp model optimization use cases and integrate debugging capabilities natively into popular third-party environments like VSCode, Cursor, and Grafana.

Analyze and understand existing data collection, analysis, and visualization workflows with deep introspection across frameworks, XLA, and the runtime stack. Work with open-source ML inference frameworks like vLLM and SGLang to identify performance improvement opportunities within Xprof. Partner with teams owning different ML stack components to understand performance optimization needs and provide insights into bottlenecks.

Qualifications

A Bachelor's degree or equivalent practical experience is required. You should possess at least 8 years of experience in software engineering, machine learning infrastructure, computer architecture, distributed computing, and people management, with a strong background in debugging tools and effective communication.

Proven experience in people management is essential, alongside a track record of building infrastructure to enhance the performance of Machine Learning (ML) systems and applications.

Preferred qualifications include expertise in Performance Optimization, GPU Programming, High-Performance Computing, and Large Language Models. Hands-on experience with ML infrastructure, frameworks like TensorFlow, JAX, PyTorch, and inference frameworks such as vLLM and SGLang is highly desirable. Contributions to open-source projects are also a plus.

Essential Skills

Software EngineeringMachine Learning InfrastructureComputer ArchitectureDistributed ComputingPeople ManagementDebugging ToolsCommunication

Good to Have

Performance OptimizationGPU ProgrammingHigh Performance ComputingLarge Language ModelsOpen Source ContributionML FrameworksGPU Performance AnalysisTPU Performance AnalysisAgentic WorkflowsTensorFlowJAXPyTorchKerasvLLMSG LangPathwaysOpen Source Software Development

Highlights

  • Actively hiring

More Details

RoleSoftware Engineering Manager, AI/ML Infrastructure and Performance Engineering
IndustryAI / Machine Learning, Cloud Computing
DepartmentSoftware Development, AI / Machine Learning
Employment TypeFull Time, Hybrid (office + remote)

About the Company

Google logo

Google

AI / Machine Learning

Software Engineering Manager, AI/ML Infrastructure and Performance Engineering at Google | SkillMX | SkillMX