Team Lead | Site Reliability Engineering | Bengaluru | Engineering | Hybrid Cloud Engineering

Deloitte

3–5 yrs Bengaluru, Coimbatore Full Time Hybrid (office + remote)
Deloitte logo
Posted : 1 week ago
Actively hiring

Job description

Join our dynamic Engineering, AI & Data team as a Team Lead, focusing on Site Reliability Engineering and DevOps. This role is instrumental in empowering mission-critical solutions by modernizing systems and implementing new technology platforms. We aim to improve financial performance, accelerate digital businesses, and fuel growth through innovation.

As a leader in SRE DevOps, you will guide the implementation of best practices, ensuring application reliability, system scalability, and operational efficiency. You will collaborate closely with development, operations, and platform teams to deliver robust cloud solutions.

Our focus is on scalable, resilient, and high-performing cloud environments using cutting-edge DevOps and SRE principles. The team excels in designing, implementing, and operating enterprise-grade systems with cloud-native technologies and modern development methodologies.

Responsibilities

Apply advanced DevOps and Site Reliability Engineering (SRE) principles to significantly enhance system availability and performance.

Develop and maintain robust CI/CD pipelines for seamless continuous integration and delivery.

Provide comprehensive support for the deployment, monitoring, and ongoing maintenance of applications across various environments.

Actively participate in incident management, conducting thorough Root Cause Analysis (RCA) and postmortems.

Foster strong collaboration with cross-functional teams to ensure smooth and efficient project delivery.

Drive automation initiatives to minimize manual efforts and maximize operational efficiency.

Engage in Agile ceremonies, including sprint planning and retrospectives, to contribute to team improvement.

Continuously strive to enhance system stability, reliability, and observability through proactive measures.

Qualifications

Demonstrate 3–5 years of progressive experience in DevOps, SRE, or Cloud Engineering.

Possess hands-on expertise with CI/CD tools such as Azure DevOps, Jenkins, or Git.

Exhibit a solid understanding of cloud platforms, specifically Azure and/or AWS.

Showcase experience with Infrastructure as Code (IaC) tools, including Terraform, ARM, or Bicep, at a working knowledge level.

Apply basic to intermediate knowledge of Kubernetes and containerization technologies.

Proficiency in scripting languages like PowerShell, Bash, or Python is essential.

Familiarity with monitoring tools such as Prometheus, Grafana, or the ELK stack is required.

Possess a strong understanding of SRE concepts, including SLI, SLO, and error budgets.

Knowledge of DevSecOps practices is necessary.

Experience with guiding the adoption of serverless and event-driven architectures using Lambda and SNS is beneficial.

Experience overseeing cloud data and analytics platforms, leveraging Databricks and LakeFlow, is advantageous.

Ability to manage cloud financials, FinOps, forecasting, chargeback/showback models, and optimization initiatives is a plus.

Capability to collaborate with InfoSec and compliance teams to meet ISO, SOC2, HIPAA, PCI-DSS requirements is important.

Possess at least 3+ years in IT infrastructure and cloud services, with a minimum of 5+ years of deep AWS experience.

Exhibit strong expertise in core AWS services, including EC2, S3, VPC, IAM, RDS, Lambda, ECS/EKS, SNS, CloudWatch, and Route 53.

Extensive experience with Terraform and enterprise-scale IaC implementations is highly valued.

Essential Skills

SRE & Reliability EngineeringSLIs (Service Level Indicators)SLOs (Service Level Objectives)Monitor application uptimelatencyand error ratesSupport incident resolution and participate in RCA discussionsContribute to improving MTTR and overall system reliabilityCI/CD & DevOpsBuild and maintain CI/CD pipelines using: Azure DevOpsJenkinsGitHub ActionsAutomate buildtestand deployment workflowsSupport release management activitiesProvision and manage cloud resources in: Azure / AWS environmentsImplement Infrastructure as Code using: Terraform / ARM / Bicep (basic to intermediate level)Support multi-environment setup (DevUATProd)Docker for containerizationKubernetes (AKS / EKS – basic to intermediate level)Support deployment of microservicesAssist in scaling and managing workloadsConfigure and use monitoring tools: Azure MonitorCloudWatchPrometheusGrafana (basic exposure)Analyze logs and alerts to identify issues proactivelyFollow DevSecOps practicesIAM / RBACSecrets Management tools (Key Vault / Secrets Manager)Support compliance with security standardsIdentify performance bottlenecks and suggest improvementsAssist in implementing: Auto-scalingResource optimizationSupport cost optimization initiativesHands-on experience with: CI/CD tools (Azure DevOps / Jenkins / Git)Understanding of: Cloud platforms (Azure and/or AWS)Experience with: Infrastructure as Code (Terraform / ARM / Bicep – working knowledge)Basic to intermediate knowledge of: Kubernetes and containersScripting skills: PowerShell / Bash / PythonMonitoring tools (PrometheusGrafanaELK)Understanding of: SRE concepts (SLISLOerror budgets)Knowledge of: DevSecOps practicesGuide adoption of serverless and event-driven architectures using Lambda and SNSOversee cloud data and analytics platforms leveraging Databricks and LakeFlowManage cloud financialsFinOpsforecastingchargeback/showback modelsand optimization initiativesCollaborate with InfoSec and compliance teams to meet ISOSOC2HIPAAPCI-DSS requirementsStrong expertise in AWS services including EC2S3VPCIAMRDSLambdaECS/EKSSNSRoute 53Extensive experience with Terraform and enterprise-scale IaC implementationsExperience with DatabricksLakeFlowand cloud data platforms

Good to Have

Certification (good to have): Azure DevOps / AWS / Kubernetes

Highlights

  • Actively hiring

More Details

RoleTeam Lead | Site Reliability Engineering | Bengaluru | Engineering | Hybrid Cloud Engineering
IndustryEngineering, AI / Machine Learning
DepartmentDevOps / Cloud
Employment TypeFull Time, Hybrid (office + remote)

About the Company

Deloitte logo

Deloitte

Engineering

Team Lead | Site Reliability Engineering | Bengaluru | Engineering | Hybrid Cloud Engineering at Deloitte | SkillMX | SkillMX