Join Google Cloud's Memorystore team as a Site Reliability Manager and play a pivotal role in making Memorystore the premier in-memory cloud database. This is an opportunity to enhance the reliability, scalability, and security of our fully managed Redis and Valkey offerings, serving a wide array of applications with sub-millisecond data access.
Our team focuses on delivering insightful service metrics, scaling alerting infrastructure, and continuously mitigating risks in production. We are driven by innovation, particularly in finding efficient ways to manage systems amidst growing AI demand. By leveraging expertise in system architecture, we aim to build automation and self-repair mechanisms to boost customer satisfaction and reduce operational overhead, while integrating common GPP services for product enhancement.
Lead a talented team of software and systems engineers, guiding iteration and task planning while fostering a supportive team environment. Set strategic team priorities and drive OKR execution in collaboration with SRE leadership and cross-functional partners.
Cultivate strong relationships with development and cross-functional teams, establishing technical credibility and influencing the team's technical direction and delivery quality. Participate actively in the on-call rotation, promoting technical curiosity, a growth mindset, and a blameless postmortem culture within the team.
A Bachelor’s degree in Computer Science, a related field, or equivalent practical experience is required. Candidates should possess a minimum of 8 years of experience in software development, with proficiency in at least one programming language.
Preferred qualifications include experience with languages such as C++, C, Java, Python, or Go, alongside expertise in engineering or operations roles for distributed systems and large-scale environments. Strong problem-solving skills and a deep understanding of global-scale distributed systems are highly valued. The ideal candidate will also have proven experience in building and managing results-oriented engineering teams and possess excellent communication skills with a strong sense of ownership.
Cloud Computing