Stop Bad Merges with an LLM Eval Gate

Hosted by Bruno Gonçalves

Wed, Oct 7, 2026

6:00 PM UTC (30 minutes)

Virtual (Zoom)

Free to join

Invite your network

Go deeper with a course

Build a Production-Grade LLM Eval Harness
Bruno Gonçalves
View syllabus

What you'll learn

How to set regression thresholds your team can defend

Thresholds come from your metric's measured spread, not a gut number. Set them once, argue never.

How to wire an eval suite into GitHub Actions

A CLI run on every pull request, results posted to the check. The setup fits in one YAML file.

How to make a failing eval block a merge on its own

The gate flips red before users see the regression. Watch a bad prompt change die in the queue.

Why this topic matters

Right now, a bad prompt change reaches production the same way a good one does: nobody measured either. Quality depends on someone remembering to check, and memory loses to deadlines every time. A CI gate makes the check automatic. The eval runs on every pull request, and a regression dies in the queue instead of in front of users. Vigilance does not scale. Policy does.

You'll learn from

Bruno Gonçalves

PhD physicist and corporate trainer.

I earned a PhD in the Physics of Complex Systems in 2008 and held a tenured faculty position at Aix-Marseille Université and served as a Data Science fellow at NYU’s Center for Data Science before moving to Industry.

Now I consult in Generative AI, Machine Learning, and Blockchain Analytics.

The teaching runs through it all. I run corporate trainings, and my published video courses cover NLP, data visualization, and time series analysis. My newsletter, Data For Science, reaches 4,000+ subscribers, and every post ships with a working notebook.

In my sessions, every claim gets a number. You leave with code that runs on Monday morning.

Previously at

Data For Science
JPMorgan Chase & Co.
TRM Labs
New York University
See all products from Bruno

Sign up to join this lesson

By continuing, you agree to Maven's Terms and Privacy Policy.