Senior Software Engineer
Microsoft
Microsoft
Join Microsoft's cutting-edge AI infrastructure team, instrumental in shaping the next generation of large-scale AI training and inference. The AI Frameworks (AIFx) Networking & Systems Tools (NeST) organization builds foundational system software powering Microsoft's Maia accelerator platforms. Our work spans pre-silicon development, hardware bring-up, cloud integration, and production cloud infrastructure, ensuring AI accelerators operate reliably and efficiently at cloud scale.
Our India Development Centre is establishing end-to-end engineering expertise across the Maia system software stack. We operate at the intersection of distributed cloud systems and low-level accelerator software. This includes control-plane services, host and device management, accelerator virtualization, Kubernetes infrastructure, Developer/Debugger tools, hardware lifecycle management, and reliability/diagnostics.
We are seeking a Senior Software Engineer to design and develop sophisticated AI Accelerator Agent software and infrastructure for managing accelerator devices at cloud scale. This role involves deep engagement with distributed systems, networking, operating systems, and device software in both pre-silicon and post-silicon environments, with a strong emphasis on high-performance communication, device enablement, reliability, and scalability.Microsoft is dedicated to empowering every person and organization globally. As employees, we foster a growth mindset, innovate to support others, and collaborate to achieve shared objectives. We uphold our values of respect, integrity, and accountability daily, cultivating an inclusive culture where everyone can excel.
Key responsibilities include designing and developing AI Accelerator Agent software for comprehensive device configuration, monitoring, and lifecycle management. You will build scalable communication channels between control-plane services, host agents, and accelerator devices. This involves developing and optimizing networking for high throughput, low latency, and resilient distributed communication.
Enable accelerator software across various development platforms, from pre-silicon (FPGA, emulation, simulation, virtual platforms) through post-silicon hardware. Develop device-enablement software that seamlessly integrates with operating systems, drivers, firmware, and accelerator hardware. Construct AI hardware infrastructure capabilities and tools for node provisioning, configuration, health monitoring, diagnostics, and lifecycle management.
Implement robust mechanisms for device health, state management, fault detection, and recovery. Collaborate effectively with cross-functional teams including cloud, networking, driver, firmware, hardware, and infrastructure groups to deliver integrated, end-to-end solutions. Apply AI-assisted engineering practices across the entire development lifecycle—architecture, design, development, testing, debugging, and documentation—to enhance engineering velocity and software quality. Leverage AI-assisted workflows to accelerate code comprehension, network performance analysis, root-cause investigation, test development, and systems debugging.
Utilize AI to automate repetitive engineering tasks and expedite the diagnosis of complex issues within distributed systems, networking, and device-level contexts. Employ AI-assisted learning to deepen expertise in distributed systems, networking, accelerator architecture, drivers, firmware, and hardware infrastructure. Develop and disseminate reusable AI-assisted engineering practices to boost team productivity, technical learning, and domain competency.
Required qualifications include a Bachelor's Degree in Computer Science, Computer Engineering, or a related technical field, or equivalent practical experience with a minimum of 8 years of industry-relevant experience. Demonstrate strong programming proficiency in C/C++, C#, Rust, Go, or comparable systems programming languages. Possess experience developing systems software, distributed systems, or device-enablement software. A solid understanding of computer networking, network protocols, distributed communication, and optimization of throughput and latency is essential.
Experience enabling software across hardware development or validation environments, encompassing both pre-silicon and/or post-silicon platforms, is required. Exhibit a strong grasp of concurrency, asynchronous programming, reliability principles, state management, and failure recovery strategies. Possess strong system-level debugging and problem-solving capabilities. Crucially, demonstrate the ability to effectively apply AI-assisted engineering tools and workflows to enhance development velocity, debugging efficiency, testing effectiveness, and overall engineering quality.
Preferred qualifications include experience developing host/device agents or AI accelerator management software. Familiarity with GPU, AI accelerator, or heterogeneous compute infrastructure is advantageous. Experience with FPGA, hardware emulation, simulation, or virtual platform environments is a plus. Experience building infrastructure or tooling for compute-node provisioning, health management, diagnostics, and device lifecycle management is highly desirable.
Further preferred experience includes high-performance networking and large-scale distributed systems. Knowledge of Linux, device drivers, firmware interfaces, PCIe, SR-IOV, PF/VF, or device virtualization is beneficial. Experience with telemetry, profiling, tracing, and performance diagnostics is valued. Experience applying AI-assisted engineering to code analysis, network/system performance investigations, debugging, testing, and engineering automation is a strong advantage. Proficiency using AI-enabled workflows to rapidly build technical expertise across complex systems and hardware domains is also preferred.
Microsoft Corporation
Computer Software