AI Researcher & Founder at ModelsLive
Getting an AI feature to work once is the easy part. You run it a few times, the outputs look reasonable, and you ship. Then the complaints start. The answers are wrong in ways you never saw in testing, the invoice climbs every month, and when you change a prompt to fix one thing, you have no way to know what else moved.
Most teams handle this by watching dashboards and reacting to whatever users report loudest. That works until it does not, and by then the fix is expensive.
This workshop closes that gap over three live sessions. You bring your own system. In the first you build an evaluation set that shows where it fails and how often. In the second you find which calls drive your cost and move the easy ones to a cheaper model. In the third you put monitoring and a rollback path in place before you need them.
You leave with three things you built yourself: an evaluation set for your own system, a cost breakdown of your own calls, and a rollback plan you can use the same week.
Taught by an AI researcher who has published 35 papers at ACL, CVPR and EMNLP, reviewed over 100 times for the major AI conferences, and ships AI products in production as the founder of ModelsLive.
Leave with an evaluation set, a cost breakdown, and a rollback plan for your own AI system, built during the workshop.
Walk through a real failure catalog from a live system, start to finish
Use a template for labeling cases that takes minutes per example
Run your own system against the set you built during the session
Learn the questions that convert "the output feels off" into something testable
Practice on real complaint text, not cleaned-up examples
Leave with a labeling rubric your team can apply without you
Understand how many examples you need before a gap means anything
Use a simple check that takes one line of code, no statistics background needed
See a worked case where an apparent gain disappeared under scrutiny
Break down a real monthly bill by call type and volume
Spot the patterns that quietly consume most of the budget
Map your own usage during the session
Learn where the line sits between cases that need your best model and cases that do not
Apply a routing pattern you can implement the same week
Verify quality held using the evaluation set from session one
Pick the few signals worth watching instead of monitoring everything
See the failure modes that only appear once real users arrive
Set up a rollback you can trigger without redeploying everything
Turn user complaints into labeled failure cases, build a 50-example set against your own system, and learn when a score gap between two versions actually means something.
Find the call patterns driving your bill, route the easy cases to a cheaper model, and verify quality held using the evaluation set you built in session one.
Work through the failures that only appear with real users, pick the few signals worth monitoring, and set up a rollback you can trigger without redeploying everything.
AI Researcher & Founder at ModelsLive
Engineers who have shipped an AI feature and are now fielding complaints they cannot reproduce or explain.
Product managers responsible for an AI feature who need a way to judge whether it is working beyond asking the team how it feels.
Founders and technical leads running AI in production who want the cost and failure rate under control before scaling further.
Live sessions
Learn directly from Yuyan (Yolanda) Chen in a real-time, interactive format.
Three sessions in one day
Three 90-minute sessions taught live over Zoom on the same day, with time for questions on your own system.
Work on your own system
You bring your own AI feature and build against it during the sessions, so you leave with real artifacts rather than examples.
Templates you keep
A labeling rubric, an evaluation set template, and a monitoring checklist you can hand to your team.
Recordings
Every session is recorded and shared afterwards, so you can revisit the walkthroughs while applying them.
Small cohort
Group size stays small enough that we can look at individual systems during the sessions.
Maven Guarantee
Your purchase is backed by the Maven Guarantee.
Maven for Teams
Reimbursement
Get your company to pay
Everything L&D needs: email template, receipts, and certificate of completion.
Get reimbursedTeam discount
Learn with your teammates
Save 20%+ when 2 or more teammates enroll in the same cohort.
Save 20%+ with a teamPrivate cohort
Run a cohort for your org
A dedicated cohort with a custom schedule and curriculum, tailored to your team.
Book a private cohort$600
USD
12–4:30pm EST