Prove Your Prompt Change Actually Helped

Hosted by Bruno Gonçalves

Wed, Sep 9, 2026

6:00 PM UTC (30 minutes)

Virtual (Zoom)

Free to join

Invite your network

Go deeper with a course

Build a Production-Grade LLM Eval Harness
Bruno Gonçalves
View syllabus

What you'll learn

How to run a paired test on two prompt versions

Same items, both prompts, one test. The pairing cancels item difficulty and isolates the change you made.

How to tell a real improvement from noise

A p-value and an effect size, read in plain language. Know when a two-point gain means nothing.

How to pick a sample size that settles the argument

Small samples flip verdicts. Learn the size where the verdict holds, before you run the comparison.

Why this topic matters

Every team ships prompt changes on opinion. Someone eyeballs five outputs, declares victory, and deploys. Half those wins are noise, and the losses surface as user complaints. A paired test settles the question in minutes with math instead of seniority. The engineer who can run one ends debates, ships with proof, and gets trusted with the next model decision.

You'll learn from

Bruno Gonçalves

PhD physicist and corporate trainer.

I earned a PhD in the Physics of Complex Systems in 2008 and held a tenured faculty position at Aix-Marseille Université and served as a Data Science fellow at NYU’s Center for Data Science before moving to Industry.

Now I consult in Generative AI, Machine Learning, and Blockchain Analytics.

The teaching runs through it all. I run corporate trainings, and my published video courses cover NLP, data visualization, and time series analysis. My newsletter, Data For Science, reaches 4,000+ subscribers, and every post ships with a working notebook.

In my sessions, every claim gets a number. You leave with code that runs on Monday morning.

Previously at

JPMorgan Chase & Co.
TRM Labs
New York University
See all products from Bruno

Sign up to join this lesson

By continuing, you agree to Maven's Terms and Privacy Policy.