Manager | ITSM | Bengaluru | Engineering | Platform Development & Integration

Deloitte

10–15 yrs Bengaluru Full Time Work from office
Deloitte logo
Posted : 1 week ago
Actively hiring

Job description

Join our Enterprise Technology & Performance team and help organizations build, operate, and optimize resilient cloud-native platforms. We are seeking a seasoned Lead Production Incident Manager to oversee enterprise production operations, incident management, and cloud infrastructure reliability within extensive AWS environments. The ideal candidate will bring deep expertise in production support, Site Reliability Engineering (SRE), cloud technologies, and ITIL-based service management, with a preference for experience in Banking & Financial Services, particularly in Cards & Payments and Mobile Applications. Enterprise technology is a critical driver of business excellence, innovation, and long-term growth.

Responsibilities

This role is pivotal in managing secure, scalable, and highly available AWS cloud infrastructure, leading enterprise production support and incident management. You will drive service reliability, production governance, automation, disaster recovery, and operational excellence for mission-critical applications.

Key responsibilities include leading enterprise-wide Incident, Problem, and Change Management following ITIL best practices. You will own the end-to-end lifecycle of production incidents (P1–P4), ensuring swift resolution, clear communication, effective escalation, and timely closure. Lead service recovery efforts to restore critical business services within agreed SLAs. Act as the central communication hub during production outages, bridging technical teams, business stakeholders, and clients. Conduct war rooms and cross-functional coordination during major incidents. Drive Root Cause Analysis (RCA), Post Incident Reviews (PIR), and implement Corrective & Preventive Actions (CAPA) to enhance service reliability. Identify recurring incidents and collaborate with engineering to reduce Mean Time To Resolve (MTTR), operational toil, and incident recurrence.

Manage 24x7 production support for enterprise applications and cloud infrastructure, including L2/L3 Infrastructure and Application Support across Linux environments. Oversee application deployments, release management, and infrastructure operations. Drive Site Reliability Engineering (SRE) initiatives focused on automation, monitoring, observability, and platform reliability. Mentor SRE and Production Support Engineers, ensuring compliance with SLAs, KPIs, SLOs, and MTTR targets. Lead Disaster Recovery (DR) planning and High Availability architecture. Manage Kubernetes and Docker container platforms for deployment and orchestration.

Monitor production environments using a suite of tools such as ELK, Kibana, Grafana, CloudWatch, Splunk, Prometheus, Nagios, Zenduty, and Site24x7. Support REST API-based applications, performing deployment validation, API verification, cache management, and health checks. Execute complex SQL queries for troubleshooting and application support. Drive cloud infrastructure optimization, automation, capacity planning, and cost optimization. Monitor customer satisfaction metrics (CSAT/NPS) and implement continuous service improvements. Collaborate with cross-functional teams to ensure highly resilient production environments.

Qualifications

We are looking for candidates with a Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related field. A strong IT background with 10–15 years of overall experience is required, including at least 8 years specifically in Enterprise Production Support, Incident Management, Cloud Infrastructure, and Site Reliability Engineering (SRE).

Demonstrated experience in the Banking & Financial Services sector, particularly with Cards & Payments and Mobile Applications, is highly advantageous. Proficiency in Incident, Problem & Change Management (ITIL) and Major Incident Management is essential. Deep knowledge of AWS cloud services, including EC2, S3, RDS, Lambda, VPC, IAM, DynamoDB, and CloudWatch, is critical. Experience in designing and managing highly available, scalable, secure, and disaster recovery-enabled cloud architectures is a must.

Hands-on expertise with Docker, Kubernetes, Jenkins, CI/CD pipelines, Infrastructure as Code (Terraform, CloudFormation, Ansible), and automation using Python or Bash is required. A solid understanding of Linux administration, networking, load balancing, SQL/NoSQL databases, REST APIs, and cloud security best practices is expected. Experience with monitoring and observability tools like Grafana, Kibana, ELK Stack, Splunk, Prometheus, Nagios, Zenduty, and Site24x7 is necessary.

Proven expertise in L2/L3 Production Support, Root Cause Analysis (RCA), Post Incident Reviews (PIR), Service Recovery, and Production Operations is vital. Experience leading 24x7 support teams, driving SLA/SLO/KPI compliance, MTTR reduction, automation initiatives, and operational excellence is key. Exceptional stakeholder management, client communication, leadership, analytical, and problem-solving skills are required to coordinate cross-functional teams during critical incidents. An ITIL Foundation Certification and AWS Certified Solutions Architect – Associate or Professional certification are preferred.

Essential Skills

Incident ManagementProblem ManagementChange ManagementITILAWSSRECloud OperationsLinux AdministrationKubernetesDockerCI/CDPythonBashSQLREST APIsMonitoring ToolsDisaster RecoveryHigh AvailabilityAutomationProduction SupportRoot Cause AnalysisStakeholder ManagementLeadership

Good to Have

ITIL Foundation CertificationAWS Certified Solutions Architect

Highlights

  • Actively hiring

More Details

RoleManager | ITSM | Bengaluru | Engineering | Platform Development & Integration
IndustryCloud Computing, Financial Services
DepartmentOperations, Cloud Engineering
Employment TypeFull Time, Work from office

About the Company

Deloitte logo

Deloitte

Cloud Computing

Manager | ITSM | Bengaluru | Engineering | Platform Development & Integration at Deloitte | SkillMX | SkillMX