Free Lesson
Build Production-Ready Agent Evals
45 min
Aug 14, 2026 11:00 AM
Virtual (Zoom)
In this video
What you'll learn
Build a representative golden dataset
Learn to create compact but effective set of normal, difficult, ambiguous and high-risk scenarios for evaluating agents
Separate deterministic checks from LLM evaluation
Understand which tests should be implemented in code and which tests may benefit from an LLM judge.
Measure business correctness, not just text quality
Learn how to evaluate whether an agent’s recommendation improves the business outcome it was designed to influence.
Create production acceptance thresholds
Define measurable release gates for:
task-completion rate;
policy compliance;
decision accuracy;
grounded responses;
Why this topic matters
A production AI agent cannot be approved on the strength of one successful demonstration.
Agents are probabilistic systems. Their outputs can vary across repeated runs, tool responses, changing context and incomplete data. This session gives participants a practical framework for answering all the evaluation related questions before deploying an agent into a live business process.
You'll learn from

Dr Ankur Narang
Dr. Ankur Narang brings 30+ yrs exp in AI, Tech across MNCs & many verticals

Kush Khurana
AI & ML leader; Venture Partner, DeepCoreX; Ashoka faculty
IBM; Apparel Group; Yatra; Oracle; Hike