Senior Team Lead | Engineering, AI & Data - Engineering | Site Reliability Engineering
Deloitte
Deloitte
Join our dynamic Engineering team as a Senior Team Lead focused on Site Reliability Engineering. You will be instrumental in managing and scaling mission-critical, production-grade distributed systems, primarily on Google Cloud Platform (GCP). This role emphasizes reliability, automation, observability, and operational excellence, aiming to achieve and maintain exceptional uptime.
Your focus will be on minimizing toil and enhancing system availability, contributing to improvements from 4 Nines to 5 Nines of uptime through robust engineering solutions. You will operate within a 24x7 environment, participating in on-call rotations to ensure the continuous operation of vital banking services.
We seek candidates with a strong aptitude for learning, a bias for action, and a deep technical understanding of cloud-native technologies. This position offers an exciting opportunity to transform technology platforms and drive business growth through innovation and engineering excellence.
Take ownership of the end-to-end reliability, availability, scalability, cost, and performance of production systems. Drive significant improvements in Mean Time To Recovery (MTTR) and Mean Time To Acknowledge (MTTA) through strategic automation and process enhancements.
Participate actively in 24x7 on-call rotations, expertly managing high-severity incidents and meticulously documenting learnings. Establish and diligently manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), Service Level Agreements (SLAs), and Error Budgets for critical services, collaborating with engineering teams to uphold SLO commitments.
Design, deploy, and manage infrastructure on Google Cloud Platform (GCP), with extensive work on GKE, Compute, networking, IAM, Load Balancers, and BigQuery. Implement and maintain infrastructure as code using Terraform, and manage containerized workloads on Kubernetes, troubleshooting issues related to pods, nodes, and networking.
Build and maintain robust CI/CD pipelines using Jenkins and GitHub, developing automation with Python and Shell scripting to reduce operational toil. Implement and manage monitoring systems using Dynatrace and Grafana, leveraging logs, metrics, and traces for deep observability to proactively identify and resolve issues. Define alerting strategies and create comprehensive runbooks.
Perform deep troubleshooting of distributed systems, microservices, and applications (Java, Golang), debugging application, infrastructure, and network-related problems. Drive continuous improvement by identifying and eliminating repetitive manual tasks and fostering a culture of reliability engineering, collaborating with development teams to enhance system design and resilience.
A Bachelor's or Master's degree is required, complemented by 4–8 years of progressive experience in Site Reliability Engineering (SRE), DevOps, or Cloud Engineering. Proven hands-on experience managing production-grade systems within 24x7 environments is essential.
Demonstrated experience with high-scale distributed systems and a solid understanding of security and compliance practices are key. Familiarity with deployment strategies such as Canary and Blue-Green is highly valued. Exposure to the banking or financial domain is considered an advantage.
Technical proficiency includes strong expertise in Google Cloud Platform (GCP) components like GKE, VPC, IAM, and Load Balancing, alongside deep expertise in Kubernetes and Docker. Strong hands-on experience with Terraform for infrastructure as code is mandatory. Experience with Jenkins and GitHub for CI/CD, coupled with strong Python and Shell scripting skills, is critical.
Candidates must have experience with observability tools such as Dynatrace and Grafana, and a working knowledge of Java and/or Golang applications with strong debugging capabilities. A strong foundation in Linux fundamentals and deep understanding of TCP/IP networking are necessary for troubleshooting network issues in distributed systems. Solid understanding of SLI, SLO, SLA, Error Budgets, and proven experience in improving MTTR and MTTA are crucial for this role.
Deloitte
Engineering