Author, AI Agent Evaluation (Springer)

You shipped an AI agent. It demos beautifully. Then a board member, a customer, or your own gut asks the question you have been avoiding: how do you know it works? Most teams answer with more testing, more prompt tuning, and a dashboard nobody trusts. The gap does not close, because the missing piece is not effort. It is a system: a written standard for what good means, an instrument that measures against it, and a loop that runs on a schedule. This is the companion course to AI Agent Evaluation (Mansoor and Bokhari, Springer 2026), cut for the person who makes the release decision rather than writes the code. Six modules take you from naming the failure you are afraid of, to writing a contract your agent has to honor, to scoring it on four dimensions, to running a loop that improves it every week. Every tool, framework and metric is introduced at concept level. After a module you can hold your own with your engineers using these terms. No lesson asks you to open a terminal. You leave with seven artifacts built on your own agent: a contract, a rubric with anchors, a benchmark set, a failure taxonomy, a release gate, a cost and latency budget, and a seven-step playbook.
You stop arguing about whether the agent is good, and start showing the number, the rubric behind it, and the decision it justifies.
Five patterns recur in production: flip-flop behavior, dependency brittleness, safety leaks, budget blowups, memory mistakes.
Each one comes with a detector that catches it before a customer does.
Three promises any stakeholder can check without reading code: do the task right, fail safely, respect the budget.
Worked end to end on a real refund-draft agent, as a table you copy line by line.
Behavior, capability, reliability and safety, each with what it actually measures and the number that should worry you.
Including how to inject chaos on purpose, so a reliability score means something.
The anchor wording carries the load, not the list of dimensions. You write anchors at three score levels and test them.
Then set numeric cutoffs for ship, ship with a flag, and hold.
Generate, score, diagnose, update, repeat, with a human approval gate so a self-optimizing agent cannot learn to break itself.
One diagnosed pattern and one fix per motion, so every gain has a traceable reason.
Cost per resolved task and 95th percentile latency, read together, turn a quality score into a release decision.
Averages hide the slow, expensive tail your customers actually notice.

Co-author of AI Agent Evaluation (Springer 2026), written for non-coders.
Founders whose agent demos beautifully, then answers the same question two ways on a Tuesday. You need the rate, not the anecdote.
Product leads who own the AI feature but not the code. Engineers report a score. You cannot tell if the product improved or the test eased.
Technical founders who stopped coding. You read a system diagram fine. You want the vocabulary back at concept level, without a code tour.
Every exercise is built on your own agent, so you leave with a contract and a rubric for something real rather than a worked example.
Everything is taught at concept level. You will hold your own with your engineers using these terms without opening a terminal.
The templates only pay off against a decision you actually have to make, so bring a release you are genuinely unsure about.

Live sessions
Learn directly from Abdullah Mansoor in a real-time, interactive format.
Lifetime access
Go back to course content and recordings whenever you need to.
Community of peers
Stay accountable and share insights with like-minded professionals.
Certificate of completion
Share your new skills with your employer or on LinkedIn.
Maven Guarantee
Your purchase is backed by the Maven Guarantee.
4 live sessions • 24 lessons
Aug
25
Live sessions
1 hr / week
Four 60-minute sessions, one a week: kickoff and diagnosis, contract clinic, rubric teardown, and the ship-gate review where you present your release decision.
Tue, Aug 25
5:00 PM—6:00 PM (UTC)
Tue, Sep 1
5:00 PM—6:00 PM (UTC)
Tue, Sep 8
5:00 PM—6:00 PM (UTC)
Tue, Sep 15
5:00 PM—6:00 PM (UTC)
Projects
1 hr / week
Seven artifacts, each built on an agent you already ship: a contract, a rubric with anchors, a benchmark set, a failure taxonomy, a release gate, a cost and latency budget, and a seven-step playbook.
Async content
1 hr / week
24 written lessons, about two hours forty-five minutes of reading in total, with diagrams you can take straight into a conversation with your engineers.
This is going to save me so much time.
Larissa Finn

From Module 1. The same billing agent, asked the same question forty times.
Maven for Teams
Reimbursement
Get your company to pay
Everything L&D needs: email template, receipts, and certificate of completion.
Get reimbursedTeam discount
Learn with your teammates
Save 20%+ when 2 or more teammates enroll in the same cohort.
Save 20%+ with a teamPrivate cohort
Run a cohort for your org
A dedicated cohort with a custom schedule and curriculum, tailored to your team.
Book a private cohort$200
USD