Agentic AI Evaluation Engineer

EY

4–7 yrs Kolkata Full Time Hybrid (office + remote)
EY logo
Posted : 2 weeks ago

Job description

EY Assurance Digital is seeking a GenAI / Agentic AI Evaluation Engineer to join their team. This role focuses on building and scaling evaluation capabilities for Generative AI, RAG-based, and Agentic AI solutions prior to deployment. It sits at the crucial intersection of AI evaluation, Responsible AI, and GenAI security.

The primary goal is to ensure these AI systems are safe, reliable, robust, and fit for purpose. You will achieve this by developing evaluation strategies, creating repeatable test harnesses, and generating auditable evidence to support deployment decisions.

Collaborate with global stakeholders, including product teams, solution architects, and risk & compliance leaders. Your responsibilities will include defining evaluation requirements, sourcing test data, conducting rigorous functional and non-functional evaluations, and proposing mitigations to reduce risk. This is a hands-on role demanding a strong AI development and evaluation mindset, with the ability to translate risk concerns into practical testing and measurable acceptance criteria.

Responsibilities

Define and implement robust evaluation strategies for GenAI systems across various use cases, including Q&A assistants, summarization, extraction, drafting, agentic systems, and multi-step workflows.

Translate business use cases into comprehensive evaluation plans, detailing scope, assumptions, success criteria, datasets, metrics, red-team scenarios, thresholds, and reporting needs.

Drive standardization by developing reusable evaluation templates, test case libraries, scoring rubrics, and reporting formats for all product teams.

Specify structured dataset requirements for product teams, ensuring thorough coverage of core user journeys, edge cases, adversarial scenarios, and bias/fairness considerations.

Build versatile evaluation pipelines for assessing answer quality, grounding and faithfulness (crucial for RAG), agentic behavior, and operational quality such as latency and cost.

Integrate LLM-as-judge methodologies with human evaluation, designing rubrics and sampling plans for calibrated assessments.

Implement automated evaluation harnesses using Python, enabling batch runs, configurable metrics, reproducible executions, and artifact storage for auditability.

Execute structured red teaming aligned with OWASP Top 10 for LLM Applications, focusing on prompt injection, data leakage, and model denial-of-service patterns.

Integrate evaluation processes into the development lifecycle, including pre-release regression gates and CI checks.

Perform adversarial testing for agentic workflows, identifying tool misuse, unauthorized actions, and potential data exfiltration.

Recommend effective mitigations such as input validation, tool sandboxing, guardrails, and policy prompting.

Produce high-quality, auditable evaluation reports, presenting clear findings, risk assessments, and control recommendations to stakeholders.

Qualifications

A strong academic foundation is essential, ideally with a bachelor's or master's degree in Data Science, Statistics, Engineering, Operational Research, or a related field with a focus on modern data architectures and processes.

Possess 4-7 years of relevant experience in areas such as ML/AI/GenAI/Agentic engineering, evaluation engineering, applied research, or security testing/red teaming.

Demonstrate strong hands-on Python proficiency for building evaluation harnesses, data processing, metric computation, and reporting pipelines.

Exhibit a practical understanding of GenAI system architectures, including RAG, embeddings, prompt orchestration, tool calling, and multi-agent systems.

Have experience in designing metrics and evaluation methods, including rubrics, automated scoring, and sampling strategies.

Familiarity with LLM risks and mitigations, particularly for enterprise contexts like data leakage, hallucinations, and prompt injection.

Possess strong security/red teaming skills, understanding OWASP Top 10 for LLM Applications and translating it into actionable test cases.

Experience with adversarial testing approaches and secure-by-design practices for LLM applications is highly valued.

Familiarity with evaluation frameworks and tooling (e.g., RAGAS, DeepEval) and experimentation practices is beneficial.

Basic DevOps practices, including Git, CI/CD, and containerization, are expected.

Excellent written communication skills are required for producing clear evaluation plans and reports for diverse audiences.

Ability to constructively challenge assumptions and influence engineering teams towards remediation is crucial.

Comfort operating in ambiguity within the fast-evolving GenAI landscape is essential.

Essential Skills

PythonGenAIAgentic AIRAGEvaluation EngineeringLLMsResponsible AISecurity TestingRed TeamingData ScienceStatisticsEngineeringOperational ResearchPrompt EngineeringAdversarial TestingCI/CDDockerGit

Good to Have

AssuranceFinanceRegulatory EnvironmentsNIST AI RMFISO/IEC 42001EU AI ActAzureProject Management

More Details

RoleAgentic AI Evaluation Engineer
DepartmentQuality Assurance
Employment TypeFull Time, Hybrid (office + remote)

About the Company

EY Global Delivery Services ( EY GDS) logo

EY Global Delivery Services ( EY GDS)

IT Consulting

Agentic AI Evaluation Engineer at EY | SkillMX | SkillMX