Site Reliability Manager

Google

5 yrs Bengaluru Full Time Hybrid (office + remote)
Google logo
Posted : today
Actively hiring

Job description

Join a leading organization at the forefront of Site Reliability Engineering (SRE), where software and systems engineering converge to build and operate robust, large-scale distributed systems. This role focuses on ensuring the reliability, uptime, and performance of critical services, with a strong emphasis on automation and infrastructure optimization.

As an SRE Manager, you will lead a team dedicated to supporting enterprise applications crucial for Google's operations, including finance, legal, supply chain, and HR. The mission is to drive service excellence through engineering innovation and a customer-centric approach, transforming Google's enterprise domain.

Responsibilities

Lead a team of 6-10 Site Reliability Engineers, providing support for Google's core enterprise services. Develop strategic roadmaps, objectives, and key results (OKRs) to enhance the maturity of managed services. Oversee the complete lifecycle of services, from initial design and deployment to ongoing operation and refinement. Support new services pre-launch through system design consultation, framework development, capacity planning, and launch reviews. Maintain live services by monitoring availability, latency, and overall system health. Drive sustainable system scaling through automation and implement changes to improve reliability and velocity. Practice effective incident response to ensure services consistently meet their objectives.

Qualifications

A Bachelor's degree in Computer Science, a related technical field, or equivalent practical experience is required. Possess a minimum of 5 years of experience in building or managing distributed systems or cloud infrastructure, with a specific focus on Kubernetes. Demonstrate 5 years of experience in people management. Familiarity with site reliability engineering principles, system design, and distributed computing is essential.

Essential Skills

KubernetesSite Reliability EngineeringSystem DesignDistributed ComputingPeople Management

Good to Have

Enterprise ToolingSAPERP

Highlights

  • Actively hiring

More Details

RoleSite Reliability Manager
DepartmentSite Reliability Engineering, Engineering Management
Employment TypeFull Time, Hybrid (office + remote)

About the Company

Google logo

Google

IT Consulting

Site Reliability Manager at Google | SkillMX | SkillMX