Free Lesson

Does Your AI Actually Work? Build the Evals That Prove It.

60 min
Jun 5, 2026 3:00 PM
Virtual (Zoom)

In this video

What you'll learn

Learn how to build evals, not a vibe check

The three moves: assert what's deterministic, judge the rest against a rubric, and calibrate it to a human.

Write a rubric that catches what matters

Good vs bad rubrics on a real example (scoring a "tell me about yourself"), plus the red-team cases that expose bias.

Leverage other models to pressure test your prompts

It's not enough to let the model that built the prompt, test the prompt. Pit Claude vs Codex vs Gemini vs Grok

Why this topic matters

You shipped the AI feature and everyone's using it. But is it actually good? Most people are checking complaints and trusting their gut. If you want to really know, you need to build evals: a rubric for what "good" means, test cases built to break it, and a judge calibrated to their own judgment. This session shows you how, live, on an example you're familiar with. Build better AI products!

You'll learn from

Will Lowrey

Will Lowrey

20 years in product leadership. Indeed, Bazaarvoice, startups. 200+ PMs coached.

Previously at

Indeed
Bazaarvoice
CGI
Intentional Product Manager
See all products from Will Lowrey