Manager | ITSM | Bengaluru | Engineering | Platform Development & Integration (Bengaluru, IN)

Deloitte

10–15 yrs Bengaluru Full Time Hybrid (office + remote)
Deloitte logo
Posted : today
Actively hiring

Job description

Spearhead enterprise-wide production operations, incident management, and cloud infrastructure reliability for large-scale AWS environments. This role is crucial for driving service resilience, governing production, and implementing automation for mission-critical applications within the Enterprise Technology & Performance team. Focus on optimizing platforms for functional excellence and enabling innovation.

Our team empowers organizations to build, operate, and enhance robust cloud-native platforms. We are seeking a seasoned Lead Production Incident Manager (IM) to guide our cloud operations, ensuring seamless incident resolution and maintaining the high availability of our AWS infrastructure. The ideal candidate brings extensive experience in production support, Site Reliability Engineering (SRE), cloud technologies, and ITIL-based service management, ideally within the Banking & Financial Services sector, particularly in Cards & Payments and Mobile Applications.

Responsibilities

Lead comprehensive Incident, Problem, and Change Management initiatives, adhering to ITIL best practices. Oversee the entire lifecycle of production incidents (P1–P4), ensuring swift resolution, clear communication, effective escalation, and thorough closure. Drive service recovery efforts to restore critical business functions within established SLAs. Serve as the central communication hub between technical teams, business stakeholders, and clients during production disruptions. Facilitate war rooms, bridge calls, and cross-functional coordination for major incidents. Champion Root Cause Analysis (RCA), Post Incident Reviews (PIR), and Corrective & Preventive Actions (CAPA) to enhance service reliability. Proactively identify recurring incidents and collaborate with engineering teams to minimize Mean Time To Repair (MTTR), reduce operational toil, and prevent future occurrences. Manage 24x7 production support for enterprise applications and cloud infrastructure, overseeing L2/L3 support across Linux environments. Drive Site Reliability Engineering (SRE) efforts focused on automation, monitoring, observability, and platform resilience. Lead disaster recovery planning and execution, including Active-Active and Active-Passive failover strategies. Manage container platforms like Kubernetes and Docker for deployment and scaling. Utilize a range of monitoring tools including ELK, Kibana, Grafana, CloudWatch, Splunk, Prometheus, Nagios, Zenduty, and Site24x7. Support REST API-based applications and validate deployments. Execute complex SQL queries for troubleshooting. Drive cloud infrastructure optimization, capacity planning, and cost reduction. Monitor customer satisfaction and implement continuous service improvements. Collaborate with various teams to ensure highly resilient production environments.

Qualifications

A Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related field is required. We seek candidates with 10-15 years of overall IT experience, including at least 8 years specifically in Enterprise Production Support, Incident Management, Cloud Infrastructure, and Site Reliability Engineering (SRE). Strong experience in the Banking & Financial Services domain, particularly with Cards & Payments and Mobile Applications, is highly valued. Expertise in Incident, Problem & Change Management (ITIL) and Major Incident Management is essential. Proficiency in AWS cloud services such as EC2, S3, RDS, Lambda, VPC, IAM, DynamoDB, and CloudWatch is a must. Experience in designing and managing highly available, scalable, secure, and disaster recovery-enabled cloud architectures is crucial. Hands-on experience with Docker, Kubernetes, Jenkins, CI/CD pipelines, Infrastructure as Code (Terraform, CloudFormation, Ansible), and automation using Python or Bash is required. A solid understanding of Linux administration, networking, load balancing, SQL/NoSQL databases, REST APIs, and cloud security best practices is necessary. Experience with monitoring and observability tools like Grafana, Kibana, ELK Stack, Splunk, Prometheus, Nagios, Zenduty, and Site24x7 is expected. Proven expertise in L2/L3 Production Support, RCA, PIR, service recovery, and production operations is critical. Experience leading 24x7 support teams, driving SLA/SLO/KPI compliance, MTTR reduction, automation, and operational excellence is required. Excellent stakeholder management, client communication, leadership, analytical, and problem-solving skills are essential for coordinating cross-functional teams during critical incidents. An ITIL Foundation Certification and/or AWS Certified Solutions Architect certification are preferred.

Essential Skills

AWSIncident ManagementProblem ManagementChange ManagementITILSite Reliability Engineering (SRE)Production SupportCloud OperationsDisaster RecoveryKubernetesDockerCI/CDInfrastructure as CodePythonBashLinux AdministrationNetworkingSQLREST APIsMonitoring ToolsELK StackKibanaGrafanaCloudWatchSplunkPrometheusNagiosZendutySite24x7

Good to Have

ITIL Foundation CertificationAWS Certified Solutions Architect

Highlights

  • Actively hiring

More Details

RoleManager | ITSM | Bengaluru | Engineering | Platform Development & Integration (Bengaluru, IN)
IndustryFinancial Services
DepartmentOperations
Employment TypeFull Time, Hybrid (office + remote)

About the Company

Deloitte logo

Deloitte

Financial Services

Manager | ITSM | Bengaluru | Engineering | Platform Development & Integration (Bengaluru, IN) at Deloitte | SkillMX | SkillMX