I build AI that has to survive an audit

Your team put a model in production and it seems to be working. Then someone above you asked what the AI spend is returning, and the honest answer was a demo and a feeling.
That is not a failure of rigor on your part. It is a structural problem with a name. On complex work, checking an answer costs about as much as producing it, so grading does not scale, so nobody grades. Ninety six percent of engineering teams use AI tools. Around twenty percent measure the result specifically.
The gap is not laziness. It is that the standard advice does not work on the problems you have. Accuracy misleads on imbalanced data. Your eval set is probably contaminated and you cannot check. And for the most valuable work your system does there is no right answer to compare against, which is where most teams give up and let a model grade a model.
There are ways out of all of that, and none is new. Metamorphic testing dates to 1998 and was never pointed at language models until recently.
This course builds one evaluation practice, for one real system of yours, across five weeks. You leave able to say what working means, prove whether you have it, and defend the number to whoever holds the budget.
Go from a demo and a feeling to an evaluation practice you can run every week and defend to whoever controls the budget.
Acceptance criteria for output that varies, and why a model that proposes a decision needs a different bar than one that assembles evidence.
The oracle problem: what you do when you genuinely cannot state the right answer.
Leave session two with a written decision statement and a first cut at what failure costs you.
Where the data comes from, who labels it, how many examples you need, and what it means when your labelers disagree.
Whether your eval data is already in the training data, how you would know, and why rephrased samples defeat naive checks.
The cold start problem, and the most common way teams fool themselves early.
Why accuracy and F1 mislead on the problems you have, and what the Matthews correlation coefficient does differently.
Setting a cut when a false positive costs more than a false negative, or the reverse.
A worked case: moving reviewer accepted output from under 30 percent to 74 percent, and what actually moved it.
Metamorphic relations: never state the correct output, only how it must move when the input changes.
Derive relations for your own system, and choose among them when you have more than you can run.
This is the session people buy the course for, and the moment ungradeable work stops being ungradeable.
Regression testing for systems whose output varies, and A/B testing across model variants.
Drift, vendor deprecation, and the migration nobody planned for.
Confidence gating and human review, with what reviewers catch feeding back into the eval set.
What actually drives inference cost, and what happens to gross margin when cost of goods is inference.
An ROI number that survives scrutiny, and what to say when the honest answer is that you do not know yet.
You present your one page version live in the final session, to a room that will argue with it.

Product & Eng Leader (AWS, VMware, HashiCorp, NGINX) · Adjunct Professor


Engineering leader who shipped it and cannot measure it.
Product leader being asked what the AI spend returns
Anyone whose team grades model output with another model
Every exercise runs against a real system of yours. Without one there is nothing to evaluate, and the coursework does not work in theory.
The exercises run on my data, but the plan is for your system. You need to know what it does, what failure costs, and who cares about it.
Every exercise is a design decision rather than an implementation. You will read code. Your engineers build what you specify.

Live sessions
Learn directly from Thomas Underhill in a real-time, interactive format.
The Lab, In Your Browser
Every exercise runs in a hosted eval harness you open in a browser. Nothing to install, nothing to configure, and no data of yours goes anywhere. You work on a dataset I built to make each week's finding land: a contaminated subset, labelers who disagree, a metric that misleads, and a judge that flips on a padded output.
Hands-on Applied Learning
You build and apply using hands-on techniques from someone who does this daily and measures real AI outcomes.
The Research Behind It, Cited
Twenty five papers mapped to the sessions they belong to, from the 1998 metamorphic testing original through contamination detection and judge fragility work from this year. Optional depth, not required reading, and it is there so you can check my work rather than take it on trust.
Your Own Evaluation Plan
You do not leave with notes. You leave with a written evaluation plan for one real system at your company: the metrics, the eval set design and its contamination risk, the thresholds and the reasoning, the metamorphic relations you will test, the review policy, the cost model, and the one page version for your executives.
A Real Case, Worked in Full
One real evaluation build, threaded through the whole course rather than dropped in as an anecdote. Reviewer accepted output improved through measurement rather than prompt tuning. You see the instrumentation, the decisions, and the two changes that did most of the work.
A Study Group
Assigned in week one and grouped by what you are evaluating rather than what industry you are in, with a specific task each week tied to the section of the plan you are already writing. Nobody should present a plan in session ten that a peer has not already argued with.
Lifetime access
Go back to course content and recordings whenever you need to.
Certificate of completion
Share your new skills with your employer or on LinkedIn.
Ten Sessions, All Recorded
Ninety minutes each, twice a week for five weeks. Thirty five minutes teaching, twenty five on a worked example against a real system, twenty applying it to yours, ten on questions. Recordings go to the cohort, though the applying is what you paid for.
Maven Guarantee
Your purchase is backed by the Maven Guarantee.
12 live sessions • 9 lessons • 5 projects
Jan
26
Session 1: Why nobody can prove it
Jan
28
Session 2: Defining the bar
Jan
29
Optional: Group office hours (optional, not recorded)
Feb
2
Session 3: Building the set
Feb
4
Session 4: Contamination
Live sessions
3 hrs / week
Tue, Jan 26
4:00 PM—5:30 PM (UTC)
Thu, Jan 28
4:00 PM—5:30 PM (UTC)
Fri, Jan 29
4:00 PM—5:30 PM (UTC)
Projects
1-3 hrs / week
Async content
1-3 hrs / week
Maven for Teams
Reimbursement
Get your company to pay
Everything L&D needs: email template, receipts, and certificate of completion.
Get reimbursedTeam discount
Learn with your teammates
Save 20%+ when 2 or more teammates enroll in the same cohort.
Save 20%+ with a teamPrivate cohort
Run a cohort for your org
A dedicated cohort with a custom schedule and curriculum, tailored to your team.
Book a private cohort$2,500
USD