Senior Team Lead | Site Reliability Engineering | Bengaluru | Engineering (Bengaluru, IN)

Deloitte

5–8 yrs Bengaluru Full Time Hybrid (office + remote)
Deloitte logo
Posted : today
Actively hiring

Job description

Join our Engineering, AI & Data team as a Senior Team Lead for Site Reliability Engineering in Bengaluru. You will be instrumental in managing and scaling mission-critical production systems on AWS, focusing on reliability, automation, and operational excellence. This role is key to enhancing system availability and driving efficiency in a 24x7 operational environment.

We seek a candidate with deep expertise in cloud-native technologies, Kubernetes, and infrastructure automation. You'll play a vital role in the complete lifecycle management of critical banking services, ensuring high levels of agility, learnability, and adaptability. An engineer with a proactive approach and a passion for rapid learning will thrive here.

Responsibilities

Take ownership of end-to-end production system reliability, availability, scalability, cost, and performance. Drive significant improvements in MTTR and MTTA through automation and enhanced incident response processes.

Participate in 24x7 on-call rotations, effectively handling high-severity incidents and documenting lessons learned. Establish and manage key operational metrics like SLIs, SLOs, SLAs, and Error Budgets, collaborating with engineering teams to uphold service level objectives.

Design, deploy, and manage infrastructure on AWS, including extensive work with Kubernetes Engine, compute, networking, IAM, load balancers, and BigQuery. Implement and manage infrastructure as code using Terraform. Deploy and manage containerized workloads, troubleshooting Kubernetes-related issues.

Build and maintain CI/CD pipelines using Jenkins and GitHub. Develop automation solutions using Python and Shell scripting to reduce operational toil. Implement and manage monitoring systems with tools like Dynatrace and Grafana, ensuring deep observability through logs, metrics, and traces. Define alerting strategies and create runbooks.

Perform in-depth troubleshooting of distributed systems, microservices, and applications. Collaborate with development teams to enhance system design and resilience. Identify and eliminate repetitive manual tasks, fostering a culture of reliability engineering.

Qualifications

A Bachelor's degree in Engineering is required, alongside 5-8 years of progressive experience in SRE, DevOps, or Cloud Engineering. Demonstrable hands-on experience managing production-grade systems in 24x7 environments is essential.

Technical expertise should include strong AWS proficiency (Kubernetes Engine, VPC, IAM, Load Balancing, KMS, etc.), deep knowledge of Kubernetes (GKE) and Docker, and robust troubleshooting skills in containerized environments. Proficiency in Terraform for infrastructure as code is a must.

Experience with CI/CD tools like Jenkins and GitHub is required, along with strong scripting skills in Python and/or Shell scripting. Familiarity with automation frameworks and observability tools such as Dynatrace/Grafana is expected. Working knowledge of Java and/or Golang applications, coupled with strong debugging skills across application and infrastructure layers, is necessary.

Solid understanding of Linux fundamentals and TCP/IP networking is crucial for debugging network issues in distributed systems. A strong grasp of SRE principles, including SLIs, SLOs, SLAs, and Error Budgets, is also required. Proven experience in improving MTTR and MTTA and managing the incident management lifecycle are key competencies.

Excellent analytical and troubleshooting skills, strong communication and stakeholder management abilities, and a proactive, ownership-driven approach are essential soft skills. Experience in high-scale distributed systems and familiarity with deployment strategies like Canary and Blue-Green are preferred.

Essential Skills

AWSKubernetesTerraformDockerJenkinsGitHubPythonShell ScriptingDynatraceGrafanaJavaGolangLinuxTCP/IPSLISLOSLAMTTRMTTAIncident ManagementCloud ArchitectureCI/CDInfrastructure as Code

Good to Have

Banking DomainSecurity PracticesCompliance PracticesCanary DeploymentsBlue-Green Deployments

Highlights

  • Actively hiring

More Details

RoleSenior Team Lead | Site Reliability Engineering | Bengaluru | Engineering (Bengaluru, IN)
DepartmentSite Reliability Engineering, Engineering
Employment TypeFull Time, Hybrid (office + remote)

About the Company

Deloitte logo

Deloitte

IT Consulting

Senior Team Lead | Site Reliability Engineering | Bengaluru | Engineering (Bengaluru, IN) at Deloitte | SkillMX | SkillMX