Software Engineering Manager, AI/ML Infrastructure and Performance Engineering

Google

8 yrs Bengaluru Full Time Hybrid (office + remote)
Google logo
Posted : today
Actively hiring

Job description

Lead the advancement of AI/ML infrastructure and performance engineering, guiding a team of talented engineers. This role is pivotal in shaping the future of machine learning systems and applications.

Leverage your deep technical expertise to provide strategic direction and management for major projects. You'll be instrumental in optimizing not just your own code, but also empowering your team to achieve peak performance.

This position involves significant contributions to product strategy and team development, working across diverse areas like artificial intelligence, natural language processing, and large-scale system design. Join us in pushing the boundaries of what's possible with software engineering at scale.

Responsibilities

Drive continuous improvements in the machine learning software/hardware stacks by providing insightful performance debugging for workloads and custom kernels. Summarize captured profile data from various sources, including trace timelines, memory usage, compiler profiles, and ML graph summaries.

Support new ML paradigms, such as horizontal scaling for upcoming TPU chips, by contributing across the end-to-end stack and analysis tools. Collaborate with ML Stack leads to understand model optimization use cases and integrate debugging capabilities into third-party environments like VSCode, Cursor, and Grafana.

Work with open-source ML inference frameworks like vLLM and SGLang to identify performance improvement opportunities within Xprof. Partner with teams managing different parts of the ML stack to understand performance optimization use cases, and assist in identifying bottlenecks within frameworks such as TorchTPU, vLLM, and SGLang.

Qualifications

Requires a Bachelor's degree or equivalent practical experience, coupled with at least 8 years of experience in software engineering. This experience must span machine learning infrastructure, computer architecture, distributed computing, and people management.

Demonstrated expertise in debugging tools and strong communication skills are essential. Significant experience in people management and building infrastructure to enhance the performance of Machine Learning (ML) systems and applications is required.

Preferred qualifications include experience with Performance Optimization, GPU Programming, High Performance Computing, and Large Language Models. Hands-on experience with ML inference frameworks and performance analysis on GPUs or TPUs is highly valued.

Essential Skills

software engineeringmachine learning infrastructurecomputer architecturedistributed computingpeople managementdebugging toolscommunication

Good to Have

Performance OptimizationGraphics Processing Unit (GPU) ProgrammingHigh Performance ComputingLarge Language ModelOpen Source ContributorGPU performance analysisTPU performance analysisagentic workflowsTensorFlowJAXPyTorchKerasvLLMSG LangPathwaysopen-source software development

Highlights

  • Actively hiring

More Details

RoleSoftware Engineering Manager, AI/ML Infrastructure and Performance Engineering
Employment TypeFull Time, Hybrid (office + remote)

About the Company

Google logo

Google

IT Consulting