DE-Cloud Platform Engineer - AIOps-GDSN02
EY
EY
Join EY and contribute to building a better working world. EY's Digital Engineering team is seeking a Cloud AIOps & Automation Engineer with 4-8 years of experience. This role focuses on designing, implementing, and operating intelligent cloud operations solutions, blending AIOps, observability, cloud engineering, automation, and Site Reliability Engineering (SRE) principles. The goal is to enhance platform reliability, minimize operational burden, and expedite incident resolution.
EY Digital Engineering is a distinctive, industry-aligned business unit offering a comprehensive suite of services that merge deep industry knowledge with robust functional capabilities and product expertise. Our practice partners with clients to analyze, shape, design, and execute digital transformation initiatives, addressing critical business challenges and opportunities in areas like strategy, customer engagement, profit optimization, innovation, and technology.
We help clients translate strategic visions into actionable technical designs and transformation plans. Through our unique combination of skills and solutions, EY's DE team empowers clients to maintain a competitive edge and profitability by developing strategies to navigate rapid changes and disruptions and by supporting the execution of complex transformations.
As a Cloud AIOps & Automation Engineer, your core responsibilities will include:
- Implementing advanced AIOps functionalities such as anomaly detection, event correlation, alert noise reduction, incident prioritization, and automated remediation. - Designing and maintaining robust observability solutions leveraging logs, metrics, traces, and events. - Building and optimizing telemetry pipelines utilizing OpenTelemetry and related technologies. - Developing sophisticated automation and self-healing workflows using languages like Python, PowerShell, Bash, or Go. - Deploying AI-powered operational assistants, streamlining incident triage, and integrating RAG-based knowledge solutions. - Seamlessly integrating observability, CI/CD, ITSM, and cloud platforms to elevate operational intelligence. - Championing SRE practices, including the definition and tracking of SLIs, SLOs, effective incident management, and driving continuous reliability improvements.
To excel in this role, you should possess:
- Proficiency in AIOps and Observability Platforms such as Dynatrace, Datadog, Splunk, New Relic, Azure Monitor, and Prometheus/Grafana. - Experience with Microsoft Azure and/or AWS cloud environments. - Strong understanding of Kubernetes and containerization technologies. - Familiarity with OpenTelemetry and telemetry pipeline concepts. - Expertise in Python for automation scripting. - Experience with CI/CD tools like GitHub Actions and/or Azure DevOps. - Solid grasp of Incident Management and Site Reliability Engineering (SRE) fundamentals.
Additionally, experience with Semantic Kernel, LangGraph, LangChain, Azure AI Foundry, Azure OpenAI, Kafka, Elasticsearch/OpenSearch, Terraform, ArgoCD, or ServiceNow integrations would be advantageous.
EY Global Delivery Services ( EY GDS)
Financial Services