Join a pioneering team focused on Site Reliability Engineering (SRE), where software and systems engineering converge to build and manage robust, large-scale, and fault-tolerant distributed systems. This role is instrumental in ensuring the reliability, uptime, and continuous improvement of critical Google Cloud services. You will tackle complex scaling challenges unique to Google Cloud, leveraging your expertise in coding, algorithms, system design, and operational excellence.
The Core Enterprise System (CES) SRE team within Corporate Engineering provides essential SRE support for Google's enterprise applications, including critical functions like Finance, Legal, Supply Chain, and HR. Our mission is to drive service excellence through engineering, innovation, and a customer-centric approach, transforming Google's enterprise domain.
Google fosters an environment of intellectual curiosity, collaborative problem-solving, and a willingness to take calculated risks in a supportive, blame-free atmosphere. We encourage self-direction on impactful projects while providing the necessary mentorship for professional growth.
Lead and manage a team of 6-10 site reliability engineers dedicated to supporting Google’s vital enterprise services.
Develop strategic roadmaps, set clear objectives, and define key results (OKRs) to enhance the maturity and effectiveness of the managed services.
Contribute to and improve the entire lifecycle of services, from initial conception and design through deployment, ongoing operation, and refinement.
Provide pre-launch support by offering system design consultation, developing essential software platforms and frameworks, performing capacity planning, and conducting thorough launch reviews.
Maintain service health post-launch by meticulously measuring and monitoring availability, latency, and overall system well-being.
Drive sustainable system scaling through advanced automation techniques and continuously evolve systems by advocating for changes that boost reliability and operational velocity.
Uphold rigorous incident response practices to ensure services consistently meet their defined service level objectives.
Possess a Bachelor's degree in Computer Science, a related technical field, or demonstrate equivalent practical experience.
Acquire a minimum of 5 years of experience in building or managing distributed systems or cloud infrastructure, with a significant focus on Kubernetes.
Accumulate at least 5 years of proven experience in people management.
Demonstrate substantial experience in site reliability engineering, system design, and distributed computing principles.
For those seeking advanced opportunities, 5 years of experience in people management, specifically managing distributed, multi-site teams through engineering managers or tech leads, is highly valued.
Experience with Enterprise tooling and technology is preferred.
Familiarity with Systems, Applications, and Products (SAP) or other Enterprise Resource Planning (ERP) systems is advantageous.
Technology