Senior Team Lead | Site Reliability Engineering | Bengaluru | Engineering
Deloitte
Deloitte
Join our dynamic Engineering team as a Senior Team Lead in Site Reliability Engineering. This role is crucial for modernizing systems and implementing cutting-edge technology products and platforms, driving financial performance and accelerating digital business growth.
We are seeking a highly skilled Site Reliability Engineer (SRE) with expertise in managing and scaling production-grade distributed systems on AWS. Your focus will be on ensuring reliability, automation, observability, and operational excellence to minimize toil and maximize system availability, aiming for an impressive 5 Nines of uptime.
This position demands deep technical knowledge in cloud-native technologies, Kubernetes, infrastructure automation, Linux administration, and TCP/IP fundamentals. Proficiency in at least one programming language and strong troubleshooting skills for distributed systems are essential. You'll be involved in the complete lifecycle management of critical banking services, operating in a 24x7 environment with on-call rotations. The role requires adaptability, a keen ability to learn, and a proactive approach to problem-solving in complex, multi-cloud distributed environments.
As a Senior Team Lead in SRE, you will own the end-to-end reliability, availability, scalability, cost, and performance of production systems. You will drive significant improvements in MTTR and MTTA through automation, runbook development, and process enhancements. This role involves participation in 24x7 on-call rotations, handling high-severity incidents, and documenting learnings.
Key responsibilities include establishing and managing SLIs, SLOs, SLAs, Error Budgets, and operational metrics, ensuring these are upheld in collaboration with engineering teams. You will partner with various engineering, operations, and cloud management teams to deliver highly reliable services promptly.
Your work will involve designing, deploying, and managing infrastructure on AWS, with extensive experience in Kubernetes Engine, compute, networking, IAM, load balancers, TLS certificates, BigQuery, Pub/Sub, and cloud logging. You'll implement and manage infrastructure using Terraform (Infrastructure as Code).
In Kubernetes and Containers, you will deploy and manage containerized workloads, troubleshoot issues related to pods, nodes, networking, and storage. You'll manage deployments using Helm and YAML, employing rollout strategies like Canary and Blue-Green.
For Automation & CI/CD, you will build and maintain pipelines using Jenkins and scripting languages (Groovy, Shell, Python), utilizing GitHub as a power user. Developing automation with Python and Shell scripting to reduce operational toil is a core function. You will also implement and manage monitoring systems using Dynatrace, Grafana, and explore logs and metrics for deep observability, defining alerting strategies and creating runbooks.
You will perform deep troubleshooting of distributed systems, microservices architectures, and applications (Java, Golang), addressing infrastructure, application, and network-related problems. Continuous improvement is vital; identify and eliminate manual tasks, drive reliability engineering practices, and collaborate with development teams to enhance system design and resilience.
We are looking for a candidate with 5-8 years of progressive experience in Site Reliability Engineering, DevOps, or Cloud Engineering, with hands-on experience managing production-grade systems in 24x7 environments.
Essential technical skills include strong expertise in AWS, Kubernetes Engine, VPC, IAM, Load Balancing, TLS Certificates, KMS, logs and metrics exploration, BigQuery, and Pub/Sub. A good understanding of cloud architecture and landing zones is expected.
You must possess strong hands-on experience with Terraform, including the ability to write and debug Terraform code from scratch. Deep expertise in Kubernetes (GKE) and Docker, along with strong troubleshooting experience in Kubernetes environments, is required.
Proficiency in CI/CD tools like Jenkins (pipeline-based) and GitHub is essential. Strong scripting skills in Python (preferred) and Shell scripting are crucial, along with experience in automation frameworks and tooling.
Experience with observability tools such as Dynatrace/Grafana, and log, metrics, and trace-based monitoring is necessary. Working knowledge of Java and/or Golang applications, coupled with strong debugging skills across application and infrastructure layers, is vital.
Solid Linux fundamentals and a deep understanding of TCP/IP networking, including the ability to debug network issues in distributed systems, are required. A strong grasp of Reliability Engineering principles like SLI, SLO, SLA, and Error Budgets is expected, along with demonstrable experience improving MTTR and MTTA, and handling the incident management lifecycle.
Soft skills include a strong analytical and troubleshooting mindset, excellent communication and stakeholder management abilities, and the capacity to work effectively in high-pressure production environments. An ownership-driven and proactive approach is highly valued.
Preferred qualifications include experience in high-scale distributed systems, exposure to the banking/financial domain, and an understanding of security and compliance practices. Experience with deployment strategies like Canary and Blue-Green is a plus.
Deloitte
Financial Services