Manager | Site Reliability Engineering | Bengaluru | Engineering | Hybrid Cloud Engineering

Deloitte

12+ yrs Bengaluru Full Time Hybrid (office + remote)
Deloitte logo
Posted : 1 week ago
Actively hiring

Job description

Join our Engineering, AI & Data team as an SRE Operations Lead in Bengaluru. This pivotal role focuses on designing and implementing a robust SRE governance framework and operational model to enhance mission-critical solutions. You will collaborate with service delivery, engineering, operations, and SRE stakeholders to assess current practices, identify areas for improvement, and create a scalable, standardized SRE operating model.

This position requires a blend of SRE technical expertise, IT operations leadership, and consulting acumen. You will translate complex operational challenges into a clear target operating model, encompassing governance, roles, service reliability, incident and problem management, operational readiness, metrics, and continuous improvement. The role also involves providing senior leadership during the transition to the new SRE-led model.

Responsibilities

Define and establish the SRE governance framework and operational model. Collaborate with stakeholders to assess current operations and identify gaps. Design a scalable and standardized SRE operating model. Translate operational challenges into a target operating model covering governance, reliability, incident/problem management, and metrics. Provide senior leadership during the transition to the proposed SRE operating model. Establish operational KPIs and reliability dashboards. Manage and coordinate distributed SRE and engineering stakeholders. Ensure operational readiness for new and existing systems.

Qualifications

A minimum of 12 years of experience in IT operations, SRE, DevOps, or platform engineering. At least 5 years of experience in SRE/DevOps/reliability leadership roles. Demonstrated success in designing SRE governance frameworks and operating models for large organizations. Proven ability to assess current operating models and translate findings into target models. Experience defining SRE governance, including SLOs, SLIs, error budgets, and incident/problem management. Strong background in major incident management, problem management, and Root Cause Analysis (RCA). Experience with centralized, federated, or hybrid SRE/Operations models. Familiarity with ITIL/ITSM processes and their integration with SRE practices. Experience with cloud platforms (AWS, Azure, GCP) and observability tools (Prometheus, Grafana, Datadog, Splunk, etc.). Experience with Kubernetes, CI/CD, and infrastructure-as-code. Experience establishing SRE Centres of Excellence is highly desirable. Strong stakeholder management skills, including engagement with senior client leadership.

Essential Skills

SREDevOpsPlatform EngineeringIT OperationsService ManagementConsultingITIL/ITSMIncident ManagementProblem ManagementRoot Cause Analysis (RCA)Service TransitionOperational ReadinessStakeholder ManagementAWSAzureGCPPrometheusGrafanaDatadogDynatraceSplunkOpenTelemetryKubernetesCI/CDInfrastructure-as-CodeAI/AIOpsAutomation

Good to Have

SRE certificationCloud certificationsDevOps certifications

Highlights

  • Actively hiring

More Details

RoleManager | Site Reliability Engineering | Bengaluru | Engineering | Hybrid Cloud Engineering
DepartmentSite Reliability Engineering, Delivery Manager, Engineering
Employment TypeFull Time, Hybrid (office + remote)

About the Company

Deloitte logo

Deloitte

IT Consulting

Manager | Site Reliability Engineering | Bengaluru | Engineering | Hybrid Cloud Engineering at Deloitte | SkillMX | SkillMX