Site Reliability Engineer II

Microsoft

1–8 yrs Hyderabad, Bengaluru, Noida Full Time Hybrid (office + remote)
Microsoft logo
Posted : today
Actively hiring

Job description

Advance AI infrastructure and power generative AI workloads as a Site Reliability Engineer on the Azure Specialized AI Infrastructure team in India. This role focuses on automating and maintaining large-scale distributed systems essential for cutting-edge AI applications and machine learning models. You will be instrumental in ensuring the reliability, scalability, and performance of AI infrastructure, guaranteeing seamless operations for mission-critical AI services. Embrace a start-up mentality, foster collaboration, and champion customer needs.

Responsibilities

Ensure the robustness, scalability, and security of AI infrastructure for HPC and AI workloads. Lead incident response, conduct thorough root cause analyses, and drive continuous improvements to minimize downtime and maximize service availability. Optimize performance by identifying and resolving bottlenecks in compute, storage, networking, and specialized hardware like GPUs and InfiniBand. Develop and maintain automation tools for deployment, monitoring, predictive analysis, and management of AI infrastructure, including containerized environments using Kubernetes and Docker. Provide expert technical guidance on cloud and AI infrastructure technologies, working with cross-functional teams to foster innovation and implement best practices.

Qualifications

A Master's Degree in Computer Science, Information Technology, or a related field with at least 1 year of technical experience in software engineering, network engineering, or systems administration. Alternatively, a Bachelor's Degree in a similar field with 6+ years of technical experience. You should also possess 8+ years of professional software engineering experience, with a minimum of 5 years dedicated to service operations, monitoring, and reliability improvement for infrastructure. Experience with incident management and reliability engineering in cloud or AI environments for at least 1 year is essential. Additional requirements include the ability to meet stringent Microsoft, customer, and government security screening requirements, including the Microsoft Cloud Background Check.

Essential Skills

Software EngineeringNetwork EngineeringSystems AdministrationService OperationsMonitoringReliability EngineeringIncident ManagementCloud EnvironmentsAI EnvironmentsAzureKubernetesDockerContainers

Good to Have

Large-scale cloud systemsDistributed systemsGPUsInfiniBandHigh-performance technologies

Highlights

  • Actively hiring

More Details

RoleSite Reliability Engineer II
DepartmentSoftware Development
Employment TypeFull Time, Hybrid (office + remote)

About the Company

Microsoft Corporation logo

Microsoft Corporation

IT Consulting

Site Reliability Engineer II at Microsoft | SkillMX | SkillMX