EY - GDS Consulting - AIA - Gen AI - Manager
EY
EY
Advance your career at EY by joining our GDS Consulting team as a Manager focused on Enterprise Observability for AI and Generative AI systems. This role offers the unique opportunity to leverage global scale, support, and an inclusive culture to foster your professional growth. You will be instrumental in building reliable, scalable, and governable AI platforms, with a specific focus on Agentic AI implementations and comprehensive observability architectures. Embrace this chance to become the best version of yourself while contributing to a better working world for all.
Lead the design, development, and optimization of enterprise-level observability solutions for AI, Generative AI, Agentic AI, and Agentic RAG systems. Establish comprehensive standards for telemetry, tracing, monitoring, and operational governance across AI solutions. Develop robust traceability frameworks to track agent reasoning, tool usage, retrieval paths, and model interactions. Build end-to-end observability for AI agents, retrieval systems, APIs, and distributed AI applications, ensuring high reliability and performance. Design and implement telemetry collection pipelines to monitor model performance, latency, cost, quality, and user interactions.
Develop monitoring and evaluation frameworks for Agentic AI workflows utilizing tools like LangGraph, LangChain, MCP integrations, and Azure AI services. Construct production-grade observability services, APIs, and monitoring components using Python and FastAPI. Create operational dashboards, alerting frameworks, and real-time monitoring solutions tailored for AI workloads. Establish processes for evaluation, benchmarking, experimentation, and continuous improvement of enterprise AI systems.
Implement strong traceability and governance controls to meet auditing, compliance, security, and Responsible AI requirements. Deploy and manage scalable AI observability solutions on Azure Kubernetes Service (AKS). Integrate telemetry data from diverse sources including AI workloads, enterprise applications, APIs, databases, and distributed systems. Design efficient data pipelines for the collection, processing, and enrichment of observability and monitoring data.
Collaborate effectively with AI engineers, platform teams, data engineers, and business stakeholders to enhance operational visibility and system reliability. Develop reusable accelerators, observability frameworks, monitoring standards, and operational best practices. Continuously stay updated on advancements in Agentic AI, observability platforms, distributed tracing, AI evaluation frameworks, and emerging technologies.
We are seeking a Manager with 8-12+ years of professional experience, including a minimum of 4 years specifically in AI engineering, observability, telemetry, distributed systems monitoring, or AI operations. A strong foundation in people, technical, or delivery leadership, with at least 2 years of experience, is essential. Proven success in delivering enterprise-scale monitoring, observability, or operational intelligence platforms is a must.
Candidates should possess deep expertise in at least three of the following areas, with direct ownership in at least two: AI observability, Agentic AI implementations, monitoring and telemetry, traceability and governance, enterprise AI platforms, distributed systems monitoring, AI evaluation frameworks, cloud-native architecture, AI operations and reliability engineering, or data engineering and operational analytics. A strong capability to translate operational and governance requirements into scalable technical solutions is critical. Comfort in collaborating across engineering, platform, business, risk, and governance stakeholders is expected.
Educational qualifications include a Bachelor's or Master's degree in Computer Science, Data Science, Engineering, or a related field. Essential technical skills encompass deep expertise in Agentic AI, Agentic AI Implementation, and Agentic RAG Systems, alongside a robust understanding of observability principles, telemetry collection, and monitoring architectures. Hands-on experience with Azure AI Platform, Azure AI services, and Azure Kubernetes Service (AKS) is required. Proficiency in Python programming and developing monitoring services with FastAPI is necessary. Familiarity with MLOps, LLMOps, AI lifecycle monitoring, Responsible AI controls, and implementing alerting frameworks is also key.
EY Global Delivery Services ( EY GDS)
IT Consulting