Free Lesson

AI Evals for Building Reliable and Consistent Products

Part of Summer of AI Building: Your AI Portfolio in 4 Steps

60 min
Jul 21, 2026 6:00 PM
Virtual (Zoom)

In this video

What you'll learn

Build a Robust AI Evaluation Framework

Move from one user to diverse data, use a 'living' Golden Set for regression, and utilize metrics beyond just accuracy

Evaluate Different Types of AI Output

Run evals for agent skills, AI products or any kind of system where a generative model is producing an output

Implement Continuous Evaluation and Avoid Pitfalls

Automate evaluation as a continuous process. Use LLMs for qualitative checks and avoid pitfalls

Why this topic matters

Most AI features fail not because of bad models, but because of bad evaluation. Personal and work-related AI systems need systematic evaluation - and AI product builders need to understand how to spec it, measure it, and improve it. You'll learn the eval loop: define "good" -> build your Golden Set -> choose your eval type -> automate & iterate.

You'll learn from

Anshumani Ruddra

Anshumani Ruddra

Product Leader and Super IC at Google

Google
Disney+ Hotstar
Zynga
See all products from Anshumani