Site Reliability Manager, Memorystore

Google

8+ yrs Bengaluru Full Time Work from office
Google logo
Posted : yesterday
Actively hiring

Job description

Join the Memorystore SRE team, dedicated to making Cloud Memorystore the most reliable, scalable, and secure in-memory cloud database globally. We provide a fully managed in-memory service offering sub-millisecond data access, high availability, and scalability for diverse applications. Our offerings include Memorystore for Redis, Memorystore for Redis Cluster, and a Valkey-based solution.

Our team excels at surfacing critical service metrics to customers and stakeholders, with a vision to scale our alerting infrastructure while maintaining manageable operational load. We rigorously verify service assumptions, mitigate risks in evolving production environments, and actively seek efficient system operation methods, especially to meet the growing demands of AI.

Leveraging deep expertise in overall system architecture, we develop automation and self-repair capabilities to enhance customer satisfaction and minimize team toil. We are committed to adopting common services and tools within Google Cloud to continuously improve our product offerings.

Responsibilities

Lead and manage a team of software and systems engineers, focusing on iteration and task planning while ensuring team well-being and manager responsibilities are met.

Define team priorities and drive the execution of OKRs in collaboration with SRE leadership and cross-functional partners.

Cultivate robust partnerships with development teams and key stakeholders across the organization.

Establish technical credibility and influence the team's technical direction and delivery of high-quality outcomes.

Participate in the primary on-call rotation, fostering technical curiosity, a growth mindset, and a blameless postmortem culture within the team.

Qualifications

Possess a Bachelor’s degree in Computer Science, a related technical field, or equivalent practical experience.

Have a minimum of 8 years of experience in software development, proficient in at least one programming language.

Demonstrated experience in an engineering or operations role focused on distributed systems and large-scale environments is highly valued.

Expertise in building and managing high-performing teams of experienced engineers to successfully deliver on large-scale projects is essential.

Proven ability to excel in problem-solving and analyzing global-scale distributed systems.

Strong communication skills and a proactive sense of ownership and drive are required.

Essential Skills

Site Reliability EngineeringSoftware DevelopmentDistributed SystemsLarge-scale EnvironmentsProblem-SolvingSystems EngineeringTeam LeadershipOncall RotationPostmortem Culture

Good to Have

C++CJavaPythonGo

Highlights

  • Actively hiring

More Details

RoleSite Reliability Manager, Memorystore
IndustryCloud Computing
DepartmentSoftware Development
Employment TypeFull Time, Work from office

About the Company

Google logo

Google

Cloud Computing

Site Reliability Manager, Memorystore at Google | SkillMX | SkillMX