Site Reliability Manager, Memorystore

Google

8+ yrs Bengaluru Full Time Hybrid (office + remote)
Google logo
Posted : today
Actively hiring

Job description

Join Google Cloud's Memorystore team as a Site Reliability Manager and play a pivotal role in making Memorystore the premier in-memory cloud database. This is an opportunity to enhance the reliability, scalability, and security of our fully managed Redis and Valkey offerings, serving a wide array of applications with sub-millisecond data access.

Our team focuses on delivering insightful service metrics, scaling alerting infrastructure, and continuously mitigating risks in production. We are driven by innovation, particularly in finding efficient ways to manage systems amidst growing AI demand. By leveraging expertise in system architecture, we aim to build automation and self-repair mechanisms to boost customer satisfaction and reduce operational overhead, while integrating common GPP services for product enhancement.

Responsibilities

Lead a talented team of software and systems engineers, guiding iteration and task planning while fostering a supportive team environment. Set strategic team priorities and drive OKR execution in collaboration with SRE leadership and cross-functional partners.

Cultivate strong relationships with development and cross-functional teams, establishing technical credibility and influencing the team's technical direction and delivery quality. Participate actively in the on-call rotation, promoting technical curiosity, a growth mindset, and a blameless postmortem culture within the team.

Qualifications

A Bachelor’s degree in Computer Science, a related field, or equivalent practical experience is required. Candidates should possess a minimum of 8 years of experience in software development, with proficiency in at least one programming language.

Preferred qualifications include experience with languages such as C++, C, Java, Python, or Go, alongside expertise in engineering or operations roles for distributed systems and large-scale environments. Strong problem-solving skills and a deep understanding of global-scale distributed systems are highly valued. The ideal candidate will also have proven experience in building and managing results-oriented engineering teams and possess excellent communication skills with a strong sense of ownership.

Essential Skills

Distributed SystemsLarge-Scale EnvironmentsProblem-SolvingTeam LeadershipProject Management

Good to Have

C++CJavaPythonGo

Highlights

  • Actively hiring

More Details

RoleSite Reliability Manager, Memorystore
IndustryCloud Computing
DepartmentSoftware Development
Employment TypeFull Time, Hybrid (office + remote)

About the Company

Google logo

Google

Cloud Computing

Site Reliability Manager, Memorystore at Google | SkillMX | SkillMX