AI Trainer & Consultant | Physics PhD

You shipped the feature. The demo impressed everyone. Then Thursday's prompt tweak broke a workflow nobody checked, and a customer found it before you did.
You know the pattern. Quality debates run on screenshots in Slack. Someone asks if the model upgrade broke anything, and the honest answer is a shrug. You tried a spreadsheet of test prompts. It went stale in a month. The vendor dashboard shows charts nobody trusts.
The problem is not effort. It is the missing harness.
In four hours you build one. You leave with a running eval harness on Inspect-AI: calibrated judges, error bars on every metric, paired tests that settle arguments, and a CLI gate that blocks bad merges in CI. A swap guide moves it onto your own product data. Monday morning, it runs.
The method is a live build. I code each piece first, we build it together, then you build it alone. I am a physicist by training, and the rule from that career applies here: every claim gets a number. The full repo, the solution branch, and the recording go home with you.
One afternoon. One harness. No more shrugs.
Build an LLM eval harness with judges, statistical tests, and a CI gate. Become the engineer who proves changes work before they ship.
Demo: I build the harness skeleton live, and you mirror it in the starter repo.
Tool: MMLU as the running dataset, split for calibration and regression.
Pattern: every lab runs on "I do, We do, You do."
Framework: exact match first, rubric scoring second, LLM-as-judge last.
Lab: score your judge against hand-labeled items until agreement holds.
Habit: never trust a judge you have not measured.
Demo: one eval, two runs, two different scores. The bootstrap explains why.
Lab: resample your own results and plot the interval.
Habit: no metric ships without its error bar.
Case study: two prompts on identical items, one paired test, verdict on screen.
Tool: the harness compare command runs the test for you.
Demo: a failing eval blocks a mock merge in GitHub Actions.
Lab: you set the thresholds on the regression split.
Capability: regressions surface before users find them.
Tool: starter and solution branches with pinned requirements.
Guide: swap the MMLU loader for your own dataset.
Capability: Monday morning, the harness runs on your product.
What a working eval harness looks like, part by part. You map the four hours ahead and orient inside the starter repo.
Load MMLU, then cut it into a calibration split for tuning judges and a regression split you never touch. Train and test discipline, applied to evals.
Exact match, rubric scoring, and LLM-as-judge, in that order. You build each layer, then calibrate the judge against hand-labeled items until agreement holds.
One eval, two runs, two different scores. The bootstrap explains the spread. You resample your results, plot the interval, and learn why single numbers lie.
Two prompts run on identical items. A paired test delivers the verdict on screen. You wire the compare command into the harness and read results like a physicist.
Thresholds on the regression split turn evals into a merge gate. A failing eval blocks a mock pull request in GitHub Actions, live.
Swap the MMLU loader for your dataset. Starter and solution branches, pinned requirements, and the recording stay yours. Monday morning, it runs on your data.

PhD physicist and corporate trainer.
AI engineers who ship LLM features on vibes. They eyeball five outputs, deploy, and want a number that ends the quality argument.
Tech leads who own quality. Asked if the model upgrade broke anything, they answer from memory. They want the gate that answers for them.
Data scientists crossing into LLM work. The statistics are home turf. They lack the harness that carries that rigor into production.
Every lab is written Python: functions, imports, and reading tracebacks. You code for most of the four hours.
You clone the starter repo, run CLI commands, and switch branches. Nothing exotic, but the terminal is home base all session.
The judging and comparison labs call a hosted model. Any major provider works. Budget $3 to $8 in API spend for the session.

Live sessions
Learn directly from Bruno Gonçalves in a real-time, interactive format.
Lifetime access
Go back to course content and recordings whenever you need to.
Community of peers
Stay accountable and share insights with like-minded professionals.
Certificate of completion
Share your new skills with your employer or on LinkedIn.
The full repo goes home with you
Starter notebooks, a solution branch, pinned requirements, and the MMLU loader.
Maven Guarantee
Your purchase is backed by the Maven Guarantee.
Maven for Teams
Reimbursement
Get your company to pay
Everything L&D needs: email template, receipts, and certificate of completion.
Get reimbursedTeam discount
Learn with your teammates
Save 20%+ when 2 or more teammates enroll in the same cohort.
Save 20%+ with a teamPrivate cohort
Run a cohort for your org
A dedicated cohort with a custom schedule and curriculum, tailored to your team.
Book a private cohort$500
USD
10am–2pm EDT