Site Reliability & Resilience -Senior manager

EY

15+ yrs Noida Full Time Hybrid (office + remote)
EY logo
Posted : today
Actively hiring

Job description

EY is seeking a seasoned Site Reliability Engineering (SRE) Senior Consultant to drive IT modernization. This role involves acting as a technical advisor to develop and implement SRE principles and frameworks across enterprise IT. The focus is on transforming IT operations to enhance resilience, achieve predictability, optimize costs, and reduce IT risk.

You will be instrumental in assessing an organization's SRE maturity and crafting strategic roadmaps to elevate it. This is an opportunity to build a career with a global leader, contributing your unique perspective to a better working world.

Responsibilities

Key responsibilities include defining Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Service Level Agreements (SLAs) for products and services. You will engineer resilient design and implementation practices throughout the product lifecycle and develop robust observability solutions to monitor performance and SLA adherence.

Efforts will be directed towards engineering out manual processes (toil) through automation, including improvements in Continuous Integration and Continuous Deployment (CI/CD) pipelines. Additionally, you will optimize IT infrastructure and operational costs (FinOps) and conduct thorough reviews to simplify and improve deployed product architectures and inter-service dependencies.

Qualifications

A minimum of 15 years of experience in software product engineering principles, processes, and systems is required. Hands-on expertise in Java/J2EE, web servers (Apache Tomcat, IBM HTTP Server), application servers (Tomcat/WebSphere), and major RDBMS like Oracle is essential.

Proficiency in CI/CD tools (Azure DevOps, GitLab CI/CD, Jenkins) and Infrastructure as Code (IaC) tools (Terraform, AWS CloudFormation, Ansible) is expected. Experience with cloud technologies (AWS, Azure, GCP) and containerization/orchestration (Docker, Kubernetes, OpenShift) is crucial, alongside familiarity with reliability and observability tools (Azure App Insight, CloudWatch, Azure Monitor, Dynatrace, Splunk, ELK Stack).

Skills in defining NFRs and SLAs/SLOs/SLIs, understanding queuing models, and Linux (RHEL) performance monitoring are necessary. Experience with Web Services, SOA, ESB (DataPower), RESTful APIs, and application design patterns, including Microservices and Spring Boot, is also vital. Proficiency in Java runtimes, JVM tuning, application server performance tuning, and troubleshooting performance/scalability/availability issues is required. Mastery of collaborative development using Git, Jira, and Confluence is expected. Knowledge of AI/ML and Data Analytics is considered a plus.

Essential Skills

Site Reliability Engineering (SRE)SRE PrinciplesService DeliveryIT Risk ManagementBusiness ResiliencePredictabilityReliabilityCost OptimizationIT InfrastructureOperationsSRE ArchitectSRE ConsultingTransformation RoadmapsIT ModernizationSRE FrameworksEnterprise ITSRE Roadmap ImplementationSRE GovernanceSRE Maturity AssessmentSLO DefinitionSLI DefinitionResilient DesignObservability SolutionsSLA AdherenceToil ReductionAutomationAutomated System ManagementCI/CD ImprovementFinOpsProduct Architecture ReviewInter-service Dependency AnalysisSimplificationJavaJ2EEApache TomcatIBM HTTP ServerWebSphereOracle RDBMSAzure DevOpsGitLab CI/CDJenkinsTerraformAWS CloudFormationAnsibleAWSAzureGCPDockerPivotalKubernetesOpenShiftAzure AppInsightCloudWatchAzure MonitorDynatraceAppDynamicsSplunkELK StackNFR DefinitionQueuing ModelsThread PoolsRequest ServicingLinux Performance MonitoringRHELWeb ServicesSOADataPowerRESTful APIsApplication Design PatternsMicroservicesSpring BootCloud Native ArchitecturesJVM TuningGarbage CollectionPerformance TuningTroubleshootingScalabilityAvailabilityThread Dump AnalysisHeap Dump AnalysisQuery TuningDatabase ArchitecturePython ScriptingGitJiraConfluence

Good to Have

AI/MLData Analytics

Highlights

  • Actively hiring

More Details

RoleSite Reliability & Resilience -Senior manager
IndustryManagement Consulting
DepartmentInformation Technology, Engineering
Employment TypeFull Time, Hybrid (office + remote)

About the Company

EY Global Delivery Services ( EY GDS) logo

EY Global Delivery Services ( EY GDS)

Management Consulting

Site Reliability & Resilience -Senior manager at EY | SkillMX | SkillMX