Proving It: AI Evaluation for Engineering Leaders

Thomas Underhill

I build AI that has to survive an audit

You shipped it. Now somebody wants a number, and you do not have one.

Your team put a model in production and it seems to be working. Then someone above you asked what the AI spend is returning, and the honest answer was a demo and a feeling.

That is not a failure of rigor on your part. It is a structural problem with a name. On complex work, checking an answer costs about as much as producing it, so grading does not scale, so nobody grades. Ninety six percent of engineering teams use AI tools. Around twenty percent measure the result specifically.

The gap is not laziness. It is that the standard advice does not work on the problems you have. Accuracy misleads on imbalanced data. Your eval set is probably contaminated and you cannot check. And for the most valuable work your system does there is no right answer to compare against, which is where most teams give up and let a model grade a model.

There are ways out of all of that, and none is new. Metamorphic testing dates to 1998 and was never pointed at language models until recently.

This course builds one evaluation practice, for one real system of yours, across five weeks. You leave able to say what working means, prove whether you have it, and defend the number to whoever holds the budget.

What you’ll learn

Go from a demo and a feeling to an evaluation practice you can run every week and defend to whoever controls the budget.

  • Acceptance criteria for output that varies, and why a model that proposes a decision needs a different bar than one that assembles evidence.

  • The oracle problem: what you do when you genuinely cannot state the right answer.

  • Leave session two with a written decision statement and a first cut at what failure costs you.

  • Where the data comes from, who labels it, how many examples you need, and what it means when your labelers disagree.

  • Whether your eval data is already in the training data, how you would know, and why rephrased samples defeat naive checks.

  • The cold start problem, and the most common way teams fool themselves early.

  • Why accuracy and F1 mislead on the problems you have, and what the Matthews correlation coefficient does differently.

  • Setting a cut when a false positive costs more than a false negative, or the reverse.

  • A worked case: moving reviewer accepted output from under 30 percent to 74 percent, and what actually moved it.

  • Metamorphic relations: never state the correct output, only how it must move when the input changes.

  • Derive relations for your own system, and choose among them when you have more than you can run.

  • This is the session people buy the course for, and the moment ungradeable work stops being ungradeable.

  • Regression testing for systems whose output varies, and A/B testing across model variants.

  • Drift, vendor deprecation, and the migration nobody planned for.

  • Confidence gating and human review, with what reviewers catch feeding back into the eval set.

  • What actually drives inference cost, and what happens to gross margin when cost of goods is inference.

  • An ROI number that survives scrutiny, and what to say when the honest answer is that you do not know yet.

  • You present your one page version live in the final session, to a room that will argue with it.

Learn directly from Thomas

Thomas Underhill

Thomas Underhill

Product & Eng Leader (AWS, VMware, HashiCorp, NGINX) · Adjunct Professor

Previously at
Amazon Web Services
HashiCorp
VMware
NGINX
University of California
See all products from Thomas

Who this course is for

  • Engineering leader who shipped it and cannot measure it.

  • Product leader being asked what the AI spend returns

  • Anyone whose team grades model output with another model

Prerequisites

  • An AI system already in production, or about to be

    Every exercise runs against a real system of yours. Without one there is nothing to evaluate, and the coursework does not work in theory.

  • A system you are accountable for, and enough context to describe it

    The exercises run on my data, but the plan is for your system. You need to know what it does, what failure costs, and who cares about it.

  • No machine learning background, and no code required

    Every exercise is a design decision rather than an implementation. You will read code. Your engineers build what you specify.

What's included

Thomas Underhill

Live sessions

Learn directly from Thomas Underhill in a real-time, interactive format.

The Lab, In Your Browser

Every exercise runs in a hosted eval harness you open in a browser. Nothing to install, nothing to configure, and no data of yours goes anywhere. You work on a dataset I built to make each week's finding land: a contaminated subset, labelers who disagree, a metric that misleads, and a judge that flips on a padded output.

Hands-on Applied Learning

You build and apply using hands-on techniques from someone who does this daily and measures real AI outcomes.

The Research Behind It, Cited

Twenty five papers mapped to the sessions they belong to, from the 1998 metamorphic testing original through contamination detection and judge fragility work from this year. Optional depth, not required reading, and it is there so you can check my work rather than take it on trust.

Your Own Evaluation Plan

You do not leave with notes. You leave with a written evaluation plan for one real system at your company: the metrics, the eval set design and its contamination risk, the thresholds and the reasoning, the metamorphic relations you will test, the review policy, the cost model, and the one page version for your executives.

A Real Case, Worked in Full

One real evaluation build, threaded through the whole course rather than dropped in as an anecdote. Reviewer accepted output improved through measurement rather than prompt tuning. You see the instrumentation, the decisions, and the two changes that did most of the work.

A Study Group

Assigned in week one and grouped by what you are evaluating rather than what industry you are in, with a specific task each week tied to the section of the plan you are already writing. Nobody should present a plan in session ten that a peer has not already argued with.

Lifetime access

Go back to course content and recordings whenever you need to.

Certificate of completion

Share your new skills with your employer or on LinkedIn.

Ten Sessions, All Recorded

Ninety minutes each, twice a week for five weeks. Thirty five minutes teaching, twenty five on a worked example against a real system, twenty applying it to yours, ten on questions. Recordings go to the cohort, though the applying is what you paid for.

Maven Guarantee

Your purchase is backed by the Maven Guarantee.

Course syllabus

12 live sessions • 9 lessons • 5 projects

Week 1

Jan 26—Jan 31

    What working means

    • Jan

      26

      Session 1: Why nobody can prove it

      Tue 1/264:00 PM—5:30 PM (UTC)
    • Jan

      28

      Session 2: Defining the bar

      Thu 1/284:00 PM—5:30 PM (UTC)
    • Jan

      29

      Optional: Group office hours (optional, not recorded)

      Fri 1/294:00 PM—5:30 PM (UTC)
      Optional
    2 more items

Week 2

Feb 1—Feb 7

    An evaluation set you can trust

    • Feb

      2

      Session 3: Building the set

      Tue 2/24:00 PM—5:30 PM (UTC)
    • Feb

      4

      Session 4: Contamination

      Thu 2/44:00 PM—5:30 PM (UTC)
    2 more items

Schedule

Live sessions

3 hrs / week

    • Tue, Jan 26

      4:00 PM—5:30 PM (UTC)

    • Thu, Jan 28

      4:00 PM—5:30 PM (UTC)

    • Fri, Jan 29

      4:00 PM—5:30 PM (UTC)

Projects

1-3 hrs / week

Async content

1-3 hrs / week

Frequently asked questions

Maven for Teams

Reimbursement

Get your company to pay

Everything L&D needs: email template, receipts, and certificate of completion.

Get reimbursed

Team discount

Learn with your teammates

Save 20%+ when 2 or more teammates enroll in the same cohort.

Save 20%+ with a team

Private cohort

Run a cohort for your org

A dedicated cohort with a custom schedule and curriculum, tailored to your team.

Book a private cohort

$2,500

USD

Jan 26Feb 25
Enroll