Senior Site Reliability Engineer

Microsoft

2–12 yrs Hyderabad Full Time Hybrid (office + remote)
Microsoft logo
Posted : today
Actively hiring

Job description

Advance the frontier of Artificial Intelligence by joining the Azure Specialized AI Infrastructure team in India. This role is pivotal in supporting high-performance infrastructure tailored for generative AI workloads. As a Senior Site Reliability Engineer, you will be instrumental in automating and maintaining sophisticated, large-scale distributed systems that power cutting-edge AI applications and machine learning models. The core of this position lies in ensuring the utmost reliability, scalability, and performance of our AI infrastructure, guaranteeing seamless operations for critical AI services. We foster a dynamic, start-up-like environment that values collaboration and a strong commitment to customer success.

Microsoft is dedicated to empowering every individual and organization globally to achieve more. Our team embodies a growth mindset, consistently innovating to uplift others and collaborating to reach our collective aspirations. Daily, we uphold our core values of respect, integrity, and accountability, cultivating an inclusive culture where everyone can flourish both professionally and personally.

Responsibilities

Ensure the robustness, scalability, and security of AI infrastructure designed for HPC and AI workloads.

Lead comprehensive incident response efforts, conduct thorough root cause analyses, and implement continuous improvements to minimize downtime and maximize service availability.

Proactively identify and resolve performance bottlenecks across compute, storage, networking, and specialized hardware like GPUs and InfiniBand to boost AI system efficiency.

Develop and maintain advanced automation tools for deployment, monitoring, predictive analysis, and management of AI infrastructure, including containerized environments such as Kubernetes and Docker.

Provide expert technical guidance on cloud and AI infrastructure technologies, fostering collaboration with cross-functional teams to drive innovation and champion best practices.

Serve as a dedicated customer advocate, prioritizing service excellence and live site reliability for all AI workloads.

Stay abreast of emerging AI infrastructure technologies and industry trends, recommending adoption strategies for maximum benefit.

Qualifications

To excel in this role, candidates should possess a Master's Degree in Computer Science, Information Technology, or a related field, coupled with at least 2 years of technical experience in software engineering, network engineering, or systems administration. Alternatively, a Bachelor's Degree in a similar field along with 8 years of relevant technical experience is acceptable, as is equivalent overall experience.

We are seeking individuals with a minimum of 12 years of professional software engineering experience, including over 8 years focused on service operations, monitoring, and enhancing infrastructure reliability. A strong background of 5+ years in developing and supporting infrastructure services for AI or cloud platforms is essential. Additionally, at least 1 year of experience in incident management and reliability engineering within cloud or AI environments is required.

Candidates must also be able to meet specific security screening requirements for this position. This includes passing the Microsoft Cloud Background Check upon hire and every two years thereafter. For those with a Doctorate Degree, 3+ years of technical experience is expected. Preferred qualifications include experience with large-scale cloud or distributed systems, familiarity with Azure, Kubernetes, Docker, and container ecosystems, and hands-on work with supercomputers, AI platforms, GPUs, or InfiniBand technologies. Relevant publications or certifications are a plus.

Essential Skills

Site Reliability EngineeringAI InfrastructureLarge-scale Distributed SystemsGenerative AIMachine LearningCloud PlatformsContainerizationKubernetesDockerIncident ManagementRoot Cause AnalysisPerformance OptimizationAutomation ToolsMonitoringPredictive AnalysisService OperationsNetwork EngineeringSystems AdministrationSoftware EngineeringHigh-Performance Computing (HPC)GPUsInfiniBand

Good to Have

PublicationsCertifications

Highlights

  • Actively hiring

More Details

RoleSenior Site Reliability Engineer
IndustryInformation Technology & Services, AI / Machine Learning
DepartmentEngineering, Information Technology
Employment TypeFull Time, Hybrid (office + remote)

About the Company

Microsoft Corporation logo

Microsoft Corporation

Information Technology & Services

Senior Site Reliability Engineer at Microsoft | SkillMX | SkillMX