Free Lesson

Build Production-Ready Agent Evals

45 min
Aug 14, 2026 11:00 AM
Virtual (Zoom)

In this video

What you'll learn

Build a representative golden dataset

Learn to create compact but effective set of normal, difficult, ambiguous and high-risk scenarios for evaluating agents

Separate deterministic checks from LLM evaluation

Understand which tests should be implemented in code and which tests may benefit from an LLM judge.

Measure business correctness, not just text quality

Learn how to evaluate whether an agent’s recommendation improves the business outcome it was designed to influence.

Create production acceptance thresholds

Define measurable release gates for: task-completion rate; policy compliance; decision accuracy; grounded responses;

Why this topic matters

A production AI agent cannot be approved on the strength of one successful demonstration. Agents are probabilistic systems. Their outputs can vary across repeated runs, tool responses, changing context and incomplete data. This session gives participants a practical framework for answering all the evaluation related questions before deploying an agent into a live business process.

You'll learn from

Dr Ankur Narang

Dr Ankur Narang

Dr. Ankur Narang brings 30+ yrs exp in AI, Tech across MNCs & many verticals

Kush Khurana

Kush Khurana

AI & ML leader; Venture Partner, DeepCoreX; Ashoka faculty

IBM; Apparel Group; Yatra; Oracle; Hike

Apparel Group
Oracle
Meta
Hike
Yatra
See all products from DeepCoreX Academy