Pietro Montaldo
Aki Wijesundara, PhD
Free Lesson

When Claude Gets It Wrong: Evals for People Who Can't Code

45 min
Oct 15, 2026 1:00 PM

By continuing, you agree to Maven's Terms and Privacy Policy.

What you'll learn

Turn a vague complaint into a testable failure

Collect the outputs that went wrong and reduce them to one specific, reproducible case anyone on the team can rerun.

Build a 20-example eval set from real work

A folder and a spreadsheet, not a codebase real inputs, outputs you'd accept, and ones you'd send back.

Prove a fix worked instead of assuming it

Run the set before and after, and see whether you solved the problem or moved it.

Why this topic matters

Vibes, spot checks, evals. Most teams never leave the first, so quality becomes an argument about taste and the loudest opinion in the room wins. Someone rewrites the prompt, someone else says it feels better, and nobody can tell whether the failure was fixed or relocated. An eval set is the smallest thing that turns that argument into evidence and for most non-technical teams it is a spreadsheet, not an engineering project.

You'll learn from

Pietro Montaldo

Pietro Montaldo

AI Educator | Top Voice on AI for Growth Systems | Professor @ESADE

Aki Wijesundara, PhD

Aki Wijesundara, PhD

AI Advisor | Educator | Google AI Accelerator Alum

Google
Meta
OpenAI
Amazon Web Services
NVIDIA
See all products from TAI Labs
Get free access