Site Reliability Manager, Memorystore

Google

8+ yrs Bengaluru Full Time Work from office
Google logo
Posted : 1 week ago
Actively hiring

Job description

Join the Memorystore SRE team and contribute to making Cloud Memorystore the world's leading in-memory Cloud database. This role is pivotal in ensuring Memorystore's reliability, scalability, durability, security, and support.

Memorystore provides a fully managed in-memory service for sub-millisecond data access, offering scalability and high availability crucial for various applications. Our offerings include Redis and Memorystore for Redis Cluster, alongside a Valkey-based solution.

Our team focuses on delivering insightful service metrics to customers and stakeholders. We aim to scale our alerting infrastructure to encompass all customers while maintaining manageable operational load. Continuous verification of service assumptions and risk reduction are key, especially with the growing demands of AI.

Google Cloud empowers organizations through digital transformation, leveraging cutting-edge technology and sustainable development tools. Businesses globally rely on Google Cloud to foster growth and address their most complex challenges.

Responsibilities

Lead a dedicated team of software and systems engineers, overseeing iteration planning and task management while ensuring the team's well-being and professional growth.

Define team priorities and drive the execution of OKRs in collaboration with SRE leadership and cross-functional partners. Foster strong alliances with development teams and stakeholders.

Establish technical authority and influence the team's technical direction, ensuring high-quality deliverables. Participate actively in the primary on-call rotation, promoting technical curiosity, a growth mindset, and a blameless postmortem culture within the team.

Qualifications

A Bachelor's degree in Computer Science, a related field, or equivalent practical experience is required.

Significant experience, a minimum of 8 years, in software development using one or more programming languages is essential.

Preferred qualifications include experience with languages such as C++, C, Java, Python, or Go. Expertise in engineering or operations for distributed systems and large-scale environments is highly valued.

Demonstrated ability to build and manage high-performing engineering teams focused on large-scale project delivery. Strong analytical skills for problem-solving within global-scale distributed systems are also crucial. Excellent communication skills, coupled with a strong sense of ownership and drive, are expected.

Essential Skills

Software DevelopmentProblem-SolvingDistributed Systems

Good to Have

C++CJavaPythonGoTeam ManagementAutomationSelf-Repair

Highlights

  • Actively hiring

More Details

RoleSite Reliability Manager, Memorystore
DepartmentEngineering Manager, Site Reliability Engineering
Employment TypeFull Time, Work from office

About the Company

Google logo

Google

IT Consulting