Site Reliability Manager

Google

5 yrs Bengaluru Full Time Hybrid (office + remote)
Google logo
Posted : yesterday
Actively hiring

Job description

Join a pioneering team focused on Site Reliability Engineering (SRE), where software and systems engineering converge to build and manage robust, large-scale, and fault-tolerant distributed systems. This role is instrumental in ensuring the reliability, uptime, and continuous improvement of critical Google Cloud services. You will tackle complex scaling challenges unique to Google Cloud, leveraging your expertise in coding, algorithms, system design, and operational excellence.

The Core Enterprise System (CES) SRE team within Corporate Engineering provides essential SRE support for Google's enterprise applications, including critical functions like Finance, Legal, Supply Chain, and HR. Our mission is to drive service excellence through engineering, innovation, and a customer-centric approach, transforming Google's enterprise domain.

Google fosters an environment of intellectual curiosity, collaborative problem-solving, and a willingness to take calculated risks in a supportive, blame-free atmosphere. We encourage self-direction on impactful projects while providing the necessary mentorship for professional growth.

Responsibilities

Lead and manage a team of 6-10 site reliability engineers dedicated to supporting Google’s vital enterprise services.

Develop strategic roadmaps, set clear objectives, and define key results (OKRs) to enhance the maturity and effectiveness of the managed services.

Contribute to and improve the entire lifecycle of services, from initial conception and design through deployment, ongoing operation, and refinement.

Provide pre-launch support by offering system design consultation, developing essential software platforms and frameworks, performing capacity planning, and conducting thorough launch reviews.

Maintain service health post-launch by meticulously measuring and monitoring availability, latency, and overall system well-being.

Drive sustainable system scaling through advanced automation techniques and continuously evolve systems by advocating for changes that boost reliability and operational velocity.

Uphold rigorous incident response practices to ensure services consistently meet their defined service level objectives.

Qualifications

Possess a Bachelor's degree in Computer Science, a related technical field, or demonstrate equivalent practical experience.

Acquire a minimum of 5 years of experience in building or managing distributed systems or cloud infrastructure, with a significant focus on Kubernetes.

Accumulate at least 5 years of proven experience in people management.

Demonstrate substantial experience in site reliability engineering, system design, and distributed computing principles.

For those seeking advanced opportunities, 5 years of experience in people management, specifically managing distributed, multi-site teams through engineering managers or tech leads, is highly valued.

Experience with Enterprise tooling and technology is preferred.

Familiarity with Systems, Applications, and Products (SAP) or other Enterprise Resource Planning (ERP) systems is advantageous.

Essential Skills

Site Reliability EngineeringDistributed SystemsCloud InfrastructureKubernetesPeople ManagementSystem DesignDistributed ComputingAutomationIncident ResponseCapacity PlanningMonitoring

Good to Have

Enterprise ToolingEnterprise Resource Planning (ERP)SystemsApplicationsand Products (SAP)

Highlights

  • Actively hiring

More Details

RoleSite Reliability Manager
IndustryTechnology
DepartmentGeneral Management
Employment TypeFull Time, Hybrid (office + remote)

About the Company

Google logo

Google

Technology

Site Reliability Manager at Google | SkillMX | SkillMX