Principal Software Engineer
Microsoft
Microsoft
Microsoft is advancing its AI infrastructure to power the next era of large-scale AI training and inference. The AI Frameworks (AIFx) Networking & Systems Tools (NeST) group is developing core system software for Microsoft’s Maia accelerator platforms. This includes pre-silicon development, hardware bring-up, cloud integration, and production cloud infrastructure, ensuring AI accelerators operate reliably and efficiently at cloud scale.
The India Development Centre is building comprehensive engineering expertise across the Maia system software stack. Our work bridges distributed cloud systems and low-level accelerator software, encompassing control-plane services, host and device management, accelerator virtualization, Kubernetes-based infrastructure, developer/debugger tools, hardware lifecycle management, reliability, telemetry, and diagnostics.
We are seeking a Principal Software Engineer with profound expertise in systems and distributed software to architect and build the software foundation for large-scale AI accelerator platforms. This role involves collaboration across cloud services, operating systems, host agents, and device interfaces. You will tackle complex challenges in hardware lifecycle management, orchestration, resource management, fault detection and recovery, virtualization, observability, and infrastructure reliability.
This is a unique opportunity to shape the architecture of foundational AI infrastructure and develop systems for a diverse and evolving ecosystem of GPUs and AI accelerators.
Microsoft is dedicated to empowering every person and organization globally. As employees, we foster a growth mindset, innovate to support others, and collaborate towards our shared objectives. Our daily actions are guided by values of respect, integrity, and accountability, cultivating an inclusive environment where everyone can excel.
Architect and develop distributed control-plane services for provisioning, orchestrating, and managing AI accelerator infrastructure. Design scalable systems for state management, health monitoring, reconciliation, and fault recovery. Build host and device management software that connects cloud infrastructure with operating systems, drivers, firmware, and accelerator devices.
Develop infrastructure for accelerator virtualization, device assignment, isolation, and resource management. Integrate accelerator infrastructure with Kubernetes and cloud-native platforms. Design robust APIs and abstractions spanning cloud services, host software, and device interfaces.
Drive continuous improvements in reliability, security, observability, diagnostics, testing, and operational readiness. Champion AI-assisted engineering practices across architecture, design, coding, testing, debugging, code reviews, and documentation to enhance engineering velocity and software quality.
Apply AI-assisted workflows to accelerate code comprehension, root-cause analysis, test development, design exploration, and engineering automation. Identify opportunities to integrate AI into engineering workflows, reducing repetitive tasks, shortening development feedback loops, and boosting developer effectiveness.
Utilize AI-assisted learning alongside engineering fundamentals to accelerate development in systems, hardware, and AI infrastructure domains. Establish reusable engineering practices and mentor engineers on the responsible and effective application of AI throughout the software development lifecycle.
Diagnose intricate system issues across distributed services, operating systems, host software, drivers, firmware, and hardware. Lead architecture and design reviews, resolve complex cross-layer system challenges, and mentor engineers. Collaborate with hardware, firmware, OS, cloud infrastructure, and AI platform teams to deliver end-to-end solutions. Mentor engineers and contribute to elevating the technical capabilities and engineering standards of the organization.
A Bachelor’s Degree in Computer Science, Computer Engineering, or a related technical field, or equivalent practical experience, is required, along with at least 12 years of relevant industry experience. Proficiency in software engineering using languages such as C++, C, C#, Rust, or Go is essential.
Demonstrated experience in designing and developing distributed systems, cloud infrastructure, or systems software is crucial. A solid understanding of concurrency, state management, asynchronous programming, and failure recovery is necessary. Experience building reliable production software, coupled with strong debugging and problem-solving skills, is expected.
Candidates must exhibit proven technical leadership across complex, multi-team engineering projects. The ability to leverage AI-assisted engineering tools and workflows, combined with sound engineering judgment, is key to improving software development effectiveness, quality, and technical problem-solving capabilities.
An aptitude for rapidly acquiring expertise in complex technical domains and contributing to the team's overall competency through mentoring, technical guidance, and knowledge sharing is highly valued. Experience with cloud control planes, resource orchestration, infrastructure management systems, Kubernetes, containers, cloud-native infrastructure, and accelerators like GPUs is preferred.
Familiarity with PCIe, SR-IOV, PF/VF, device virtualization, host/device agents, Linux systems software, device drivers, firmware interfaces, hardware lifecycle management, health monitoring, fault recovery, telemetry, and observability in large-scale production systems is advantageous. Experience applying AI-assisted software engineering and developing AI-enabled engineering workflows to boost developer productivity and reduce repetitive effort is also preferred.
Microsoft Corporation
AI / Machine Learning