Free Lesson
Does Your AI Actually Work? Build the Evals That Prove It.
60 min
Jun 5, 2026 3:00 PM
Virtual (Zoom)
In this video
What you'll learn
Learn how to build evals, not a vibe check
The three moves: assert what's deterministic, judge the rest against a rubric, and calibrate it to a human.
Write a rubric that catches what matters
Good vs bad rubrics on a real example (scoring a "tell me about yourself"), plus the red-team cases that expose bias.
Leverage other models to pressure test your prompts
It's not enough to let the model that built the prompt, test the prompt. Pit Claude vs Codex vs Gemini vs Grok
Why this topic matters
You shipped the AI feature and everyone's using it. But is it actually good? Most people are checking complaints and trusting their gut.
If you want to really know, you need to build evals: a rubric for what "good" means, test cases built to break it, and a judge calibrated to their own judgment. This session shows you how, live, on an example you're familiar with.
Build better AI products!
You'll learn from

Will Lowrey
20 years in product leadership. Indeed, Bazaarvoice, startups. 200+ PMs coached.
Previously at
.png&w=1536&q=75)