Join our Site Reliability Engineering (SRE) team, where software and systems engineering converge to build and maintain robust, large-scale, and fault-tolerant distributed systems. As an SRE Manager, you will ensure the reliability, uptime, and rapid improvement of critical Google Cloud services. This role involves optimizing existing systems, developing infrastructure, and driving efficiency through automation, tackling unique scaling challenges inherent to Google Cloud.
Lead a team of 6-10 site reliability engineers dedicated to supporting Google's enterprise services. Develop strategic roadmaps and define objectives and key results (OKRs) to enhance service maturity. Oversee the entire service lifecycle, from design and deployment to operation and ongoing refinement. Provide pre-launch support through system design consultation, platform development, capacity planning, and launch reviews. Post-launch, focus on maintaining services by monitoring availability, latency, and overall system health, while driving sustainable scaling through automation and reliability improvements. Practice effective incident response to meet service level objectives.
A Bachelor's degree in Computer Science, a related technical field, or equivalent practical experience is required. Possess a minimum of 5 years of experience in building or managing distributed systems or cloud infrastructure, with a strong emphasis on Kubernetes. You should also have 5 years of experience in people management. Demonstrated experience in site reliability engineering, system design, and distributed computing is essential. Preferred qualifications include extensive experience in people management, particularly with distributed, multi-site teams, and familiarity with Enterprise tooling, SAP, or other ERP systems.
IT Consulting