Consultant | Site Reliability Engineering | Bengaluru | Engineering | Hybrid Cloud Engineering
Deloitte
Deloitte
This role seeks a seasoned Site Reliability Engineer (SRE) to expertly manage and scale critical, production-grade distributed systems on Google Cloud Platform (GCP). The focus is on ensuring reliability, implementing robust automation, enhancing observability, and achieving operational excellence, all while minimizing manual effort and maximizing system availability. A key objective is elevating system uptime from 4 Nines to 5 Nines through dedicated engineering initiatives.
This position demands profound technical acumen in cloud-native technologies, Kubernetes, infrastructure automation, Linux administration, and fundamental TCP/IP concepts. Proficiency in at least one programming language and strong diagnostic skills for complex distributed systems are essential. The engineer will be integral to the complete lifecycle management of mission-critical banking services, operating in a 24x7 environment with a rotational on-call schedule.
Take ownership of end-to-end production system reliability, availability, scalability, cost, and performance. Drive significant improvements in Mean Time To Recovery (MTTR) and Mean Time To Acknowledge (MTTA) by leveraging automation, enhancing runbooks, and refining processes.
Actively participate in 24x7 on-call rotations, adeptly handling high-severity incidents and meticulously documenting all learnings. Establish and diligently manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), Service Level Agreements (SLAs), Error Budgets, and critical operational metrics for essential services, collaborating closely with engineering teams to ensure SLO adherence.
Design, deploy, and manage infrastructure on Google Cloud Platform (GCP), including GKE, Compute, Networking, IAM, Load Balancers, TLS Certificates, BigQuery, Pub/Sub, and enhancing cloud logging, metrics, and log analysis. Implement and maintain infrastructure using Terraform for Infrastructure as Code.
Deploy and manage containerized applications via Kubernetes (GKE). Troubleshoot persistent issues related to pods, nodes, networking, and storage. Manage deployments using Helm and YAML, employing strategies like Canary or Blue-Green. Build and maintain CI/CD pipelines using Jenkins and possess strong GitHub proficiency. Develop automation using Python and Shell scripting to reduce operational toil.
Implement and manage monitoring solutions with Dynatrace and Grafana, analyzing logs, metrics, and traces for deep observability. Proactively identify trends and address potential problems. Define alerting strategies aligned with system behavior and SLOs, creating comprehensive runbooks. Collaborate with operations and engineering teams to resolve production incidents, isolating application and infrastructure issues and establishing effective tooling for debugging.
Perform in-depth troubleshooting of distributed systems, microservices on containerized platforms, and Java/Golang applications. Debug application, infrastructure, and network-related problems. Plan and execute continuous improvement initiatives, eliminating repetitive manual tasks and championing reliability engineering practices (DRY principle). Enhance system design and resilience in partnership with development teams.
A minimum of 4 to 8 years of relevant, progressive experience in Site Reliability Engineering (SRE), DevOps, or Cloud Engineering is required. Demonstrable hands-on experience managing production-grade systems in 24x7 operational environments is essential.
Core technical expertise should include Google Cloud Platform (GCP) services such as GKE, VPC, IAM, Load Balancing, TLS Certificates, KMS, and log/metric exploration, alongside BigQuery and Pub/Sub. A solid grasp of cloud architecture and landing zones is expected. Strong proficiency in Infrastructure as Code, particularly with Terraform, including the ability to write and debug code from scratch, is crucial.
Deep expertise in Containers and Orchestration, specifically Kubernetes (GKE) and Docker, with proven troubleshooting skills in Kubernetes environments. Hands-on experience with CI/CD and Automation tools like Jenkins (pipeline-based) and GitHub is necessary. Strong scripting capabilities in Python (preferred) and Shell scripting, along with experience in automation frameworks, are vital.
Familiarity with Observability tools such as Dynatrace/Grafana, and expertise in log, metric, and trace-based monitoring. Working knowledge of Java and/or Golang applications, coupled with strong debugging skills across application and infrastructure layers, is required. Robust Linux fundamentals and a deep understanding of TCP/IP networking, including the ability to debug network issues in distributed systems, are fundamental.
Essential reliability engineering skills include a solid understanding of SLI, SLO, SLA, and Error Budgets. Proven experience in improving MTTR and MTTA, along with demonstrated capability in managing the incident management lifecycle, is expected. Soft skills should encompass a strong analytical and troubleshooting mindset, excellent communication and stakeholder management abilities, comfort in high-pressure production settings, and a proactive, ownership-driven approach. Experience with deployment strategies like Canary and Blue-Green is considered a valuable asset.
Deloitte
Engineering