Senior Software Engineer
Microsoft
Microsoft
Join Microsoft's AI infrastructure team, driving innovation in large-scale AI training and inference. You'll contribute to foundational system software for Maia accelerator platforms, spanning pre-silicon development to cloud integration. This role focuses on building end-to-end engineering competency for the Maia system software stack, bridging distributed cloud systems and low-level accelerator software.
We are seeking a Senior Software Engineer to design and implement robust AKS/Kubernetes infrastructure for GPU and AI accelerator platforms. Your work will be critical for cluster lifecycle management, node and resource orchestration, deployment strategies, availability, observability, and the automation required for efficient cloud-scale accelerator operations. Microsoft is dedicated to empowering every individual and organization, fostering a culture of growth, innovation, and collaboration.
Design and develop scalable AKS/Kubernetes infrastructure tailored for GPU and AI accelerator environments. Build essential services and automation for cluster provisioning, configuration, seamless upgrades, and comprehensive lifecycle management. Develop sophisticated solutions for node lifecycle management, including health monitoring, proactive failure detection, and automated recovery mechanisms.
Enhance infrastructure scalability, reliability, and availability across extensive accelerator fleets. Create robust Kubernetes integrations for accelerator discovery, efficient resource management, intelligent scheduling, and seamless workload enablement. Develop dependable software and automation for the deployment, configuration, and ongoing management of accelerator infrastructure.
Improve telemetry, observability, diagnostics, and the overall operational readiness of distributed infrastructure. Tackle and resolve complex issues spanning Kubernetes, containers, Linux, networking, and accelerator hardware. Collaborate effectively with Control Plane, systems software, hardware, and platform teams to deliver integrated, end-to-end solutions.
Contribute to architecture and design reviews, uphold engineering best practices, ensure high code quality, and mentor junior engineers. Apply AI-assisted engineering practices across design, coding, testing, debugging, code reviews, and documentation to elevate engineering velocity and quality. Utilize AI-assisted workflows to accelerate code comprehension, troubleshooting, root-cause analysis, test development, and infrastructure automation.
A Bachelor's Degree in Computer Science, Computer Engineering, or a related technical field, or equivalent practical experience, coupled with 8+ years of relevant industry experience, is required. Proficiency in software development using languages such as Go, C++, C#, or Python is essential. Demonstrable experience in building distributed systems, cloud infrastructure, or platform services is crucial.
Hands-on experience with Kubernetes, containerization technologies, and cloud-native principles is a must. A strong grasp of distributed systems concepts, including scalability, concurrency, state management, resiliency, and failure recovery, is expected. Proven ability to develop reliable production software and debug intricate distributed systems is necessary, alongside strong design, problem-solving, and cross-team collaboration skills.
Experience leveraging AI-assisted software engineering tools and workflows to enhance development effectiveness, debugging, automation, and software quality is highly valued. The ability to rapidly acquire expertise in complex infrastructure technologies and apply this knowledge to real-world production engineering challenges is key.
Preferred qualifications include experience with Azure Kubernetes Service (AKS) or large-scale Kubernetes environments, cluster/node provisioning, Kubernetes controllers/operators, scheduling, or resource management. Experience operating and scaling Kubernetes-based production infrastructure, coupled with familiarity in GPU, AI accelerator, or heterogeneous compute infrastructure, is advantageous. Knowledge of Linux, containers, networking, host-level system software, accelerator virtualization, and resource isolation concepts is beneficial.
Microsoft Corporation
Information Technology & Services