Senior Site Reliability Engineer
Microsoft
Microsoft
Advance the frontier of Artificial Intelligence by joining the Azure Specialized AI Infrastructure team in India. This role is pivotal in supporting high-performance infrastructure tailored for generative AI workloads. As a Senior Site Reliability Engineer, you will be instrumental in automating and maintaining sophisticated, large-scale distributed systems that power cutting-edge AI applications and machine learning models. The core of this position lies in ensuring the utmost reliability, scalability, and performance of our AI infrastructure, guaranteeing seamless operations for critical AI services. We foster a dynamic, start-up-like environment that values collaboration and a strong commitment to customer success.
Microsoft is dedicated to empowering every individual and organization globally to achieve more. Our team embodies a growth mindset, consistently innovating to uplift others and collaborating to reach our collective aspirations. Daily, we uphold our core values of respect, integrity, and accountability, cultivating an inclusive culture where everyone can flourish both professionally and personally.
Ensure the robustness, scalability, and security of AI infrastructure designed for HPC and AI workloads.
Lead comprehensive incident response efforts, conduct thorough root cause analyses, and implement continuous improvements to minimize downtime and maximize service availability.
Proactively identify and resolve performance bottlenecks across compute, storage, networking, and specialized hardware like GPUs and InfiniBand to boost AI system efficiency.
Develop and maintain advanced automation tools for deployment, monitoring, predictive analysis, and management of AI infrastructure, including containerized environments such as Kubernetes and Docker.
Provide expert technical guidance on cloud and AI infrastructure technologies, fostering collaboration with cross-functional teams to drive innovation and champion best practices.
Serve as a dedicated customer advocate, prioritizing service excellence and live site reliability for all AI workloads.
Stay abreast of emerging AI infrastructure technologies and industry trends, recommending adoption strategies for maximum benefit.
To excel in this role, candidates should possess a Master's Degree in Computer Science, Information Technology, or a related field, coupled with at least 2 years of technical experience in software engineering, network engineering, or systems administration. Alternatively, a Bachelor's Degree in a similar field along with 8 years of relevant technical experience is acceptable, as is equivalent overall experience.
We are seeking individuals with a minimum of 12 years of professional software engineering experience, including over 8 years focused on service operations, monitoring, and enhancing infrastructure reliability. A strong background of 5+ years in developing and supporting infrastructure services for AI or cloud platforms is essential. Additionally, at least 1 year of experience in incident management and reliability engineering within cloud or AI environments is required.
Candidates must also be able to meet specific security screening requirements for this position. This includes passing the Microsoft Cloud Background Check upon hire and every two years thereafter. For those with a Doctorate Degree, 3+ years of technical experience is expected. Preferred qualifications include experience with large-scale cloud or distributed systems, familiarity with Azure, Kubernetes, Docker, and container ecosystems, and hands-on work with supercomputers, AI platforms, GPUs, or InfiniBand technologies. Relevant publications or certifications are a plus.
Microsoft Corporation
Information Technology & Services