Consultant | Site Reliability Engineering | Bengaluru | Engineering | Hybrid Cloud Engineering

Deloitte

4–6 yrs Bengaluru Full Time Hybrid (office + remote)
Deloitte logo
Posted : yesterday
Actively hiring

Job description

Join our dynamic Engineering, AI & Data team as a Site Reliability Engineering Consultant in Bengaluru. This role focuses on leading and scaling critical production systems within a hybrid cloud environment, primarily on GCP. You will be instrumental in enhancing system reliability, aiming to achieve five nines of uptime through innovative engineering solutions and operational excellence.

This position offers a unique opportunity to impact mission-critical banking services, driving improvements in availability and performance. You'll work with cutting-edge cloud-native technologies and play a key role in shaping the future of our infrastructure.

Responsibilities

Lead and mentor a team of DevOps/SRE engineers to achieve ambitious uptime goals (4 Nines to 5 Nines). Drive operational excellence by reducing Mean Time To Recover (MTTR) and Mean Time To Acknowledge (MTTA) through automation and process improvements. Establish and manage Service Level Indicators (SLI), Service Level Objectives (SLO), Service Level Agreements (SLA), and Error Budgets, ensuring accountability for upholding SLOs. Collaborate with cross-functional teams to deliver highly reliable services efficiently. Design, deploy, and manage infrastructure on GCP, including Kubernetes Engine, Compute, networking, IAM, and BigQuery. Implement Infrastructure as Code (IaC) using Terraform and orchestrate containerized workloads with Kubernetes (GKE). Develop and maintain CI/CD pipelines using Jenkins and manage deployments with Helm and YAML. Utilize scripting in Python and Shell for automation and reduce operational toil. Implement and manage monitoring systems using Dynatrace, Grafana, and explore logs and metrics for deep observability. Proactively identify and address production incidents, owning the resolution process.

Qualifications

A minimum of 4-6 years of progressive experience in SRE, DevOps, or Cloud Engineering is essential. Proven experience leading or mentoring technical teams is required. Hands-on expertise in managing production-grade distributed systems is a must. Deep proficiency in GCP, including Kubernetes Engine, VPC, IAM, Load Balancing, KMS, BigQuery, and Pub/Sub. Strong hands-on experience with Terraform for Infrastructure as Code, including writing and debugging code from scratch. In-depth knowledge of Kubernetes (GKE) and Docker, with robust troubleshooting skills. Extensive experience with Jenkins and GitHub for CI/CD practices. Skilled in Python and Shell scripting for automation purposes. Experience with observability tools like Dynatrace/Grafana, and knowledge of log, metric, and trace-based monitoring. Familiarity with Java and/or Golang applications. A solid understanding of Linux fundamentals and TCP/IP networking is crucial. Demonstrable experience improving MTTR and MTTA, along with handling the incident management lifecycle. A strong analytical and troubleshooting mindset, coupled with excellent communication and stakeholder management skills, is vital for success in this high-pressure production environment.

Essential Skills

GCPKubernetesTerraformCI/CDJenkinsGitHubPythonShell ScriptingDynatraceGrafanaSLISLOSLAError BudgetsMTTRMTTAIncident ManagementLinux AdministrationContainerizationDockerHelmJavaGolangTCP/IP Networking

Good to Have

Banking DomainSecurityComplianceCanary DeploymentsBlue-Green Deployments

Highlights

  • Actively hiring

More Details

RoleConsultant | Site Reliability Engineering | Bengaluru | Engineering | Hybrid Cloud Engineering
IndustryFinancial Services
DepartmentDevOps / Cloud
Employment TypeFull Time, Hybrid (office + remote)

About the Company

Deloitte logo

Deloitte

Financial Services

Consultant | Site Reliability Engineering | Bengaluru | Engineering | Hybrid Cloud Engineering at Deloitte | SkillMX | SkillMX