Make Your AI System Production Ready

Yuyan (Yolanda) Chen

AI Researcher & Founder at ModelsLive

Your AI feature works in a demo. Real users are a different story.

Getting an AI feature to work once is the easy part. You run it a few times, the outputs look reasonable, and you ship. Then the complaints start. The answers are wrong in ways you never saw in testing, the invoice climbs every month, and when you change a prompt to fix one thing, you have no way to know what else moved.

Most teams handle this by watching dashboards and reacting to whatever users report loudest. That works until it does not, and by then the fix is expensive.

This workshop closes that gap over three live sessions. You bring your own system. In the first you build an evaluation set that shows where it fails and how often. In the second you find which calls drive your cost and move the easy ones to a cheaper model. In the third you put monitoring and a rollback path in place before you need them.

You leave with three things you built yourself: an evaluation set for your own system, a cost breakdown of your own calls, and a rollback plan you can use the same week.

Taught by an AI researcher who has published 35 papers at ACL, CVPR and EMNLP, reviewed over 100 times for the major AI conferences, and ships AI products in production as the founder of ModelsLive.

What you’ll learn

Leave with an evaluation set, a cost breakdown, and a rollback plan for your own AI system, built during the workshop.

  • Walk through a real failure catalog from a live system, start to finish

  • Use a template for labeling cases that takes minutes per example

  • Run your own system against the set you built during the session

  • Learn the questions that convert "the output feels off" into something testable

  • Practice on real complaint text, not cleaned-up examples

  • Leave with a labeling rubric your team can apply without you

  • Understand how many examples you need before a gap means anything

  • Use a simple check that takes one line of code, no statistics background needed

  • See a worked case where an apparent gain disappeared under scrutiny

  • Break down a real monthly bill by call type and volume

  • Spot the patterns that quietly consume most of the budget

  • Map your own usage during the session

  • Learn where the line sits between cases that need your best model and cases that do not

  • Apply a routing pattern you can implement the same week

  • Verify quality held using the evaluation set from session one

  • Pick the few signals worth watching instead of monitoring everything

  • See the failure modes that only appear once real users arrive

  • Set up a rollback you can trigger without redeploying everything

Workshop agenda

  • Build an evaluation set for your own system

    Turn user complaints into labeled failure cases, build a 50-example set against your own system, and learn when a score gap between two versions actually means something.

  • Cut inference cost without losing quality

    Find the call patterns driving your bill, route the easy cases to a cheaper model, and verify quality held using the evaluation set you built in session one.

  • Ship, monitor and roll back safely

    Work through the failures that only appear with real users, pick the few signals worth monitoring, and set up a rollback you can trigger without redeploying everything.

Learn directly from Yuyan

Yuyan (Yolanda) Chen

Yuyan (Yolanda) Chen

AI Researcher & Founder at ModelsLive

See all products from Yolanda

Who this workshop is for

  • Engineers who have shipped an AI feature and are now fielding complaints they cannot reproduce or explain.

  • Product managers responsible for an AI feature who need a way to judge whether it is working beyond asking the team how it feels.

  • Founders and technical leads running AI in production who want the cost and failure rate under control before scaling further.

What's included

Yuyan (Yolanda) Chen

Live sessions

Learn directly from Yuyan (Yolanda) Chen in a real-time, interactive format.

Three sessions in one day

Three 90-minute sessions taught live over Zoom on the same day, with time for questions on your own system.

Work on your own system

You bring your own AI feature and build against it during the sessions, so you leave with real artifacts rather than examples.

Templates you keep

A labeling rubric, an evaluation set template, and a monitoring checklist you can hand to your team.

Recordings

Every session is recorded and shared afterwards, so you can revisit the walkthroughs while applying them.

Small cohort

Group size stays small enough that we can look at individual systems during the sessions.

Maven Guarantee

Your purchase is backed by the Maven Guarantee.

Frequently asked questions

Maven for Teams

Reimbursement

Get your company to pay

Everything L&D needs: email template, receipts, and certificate of completion.

Get reimbursed

Team discount

Learn with your teammates

Save 20%+ when 2 or more teammates enroll in the same cohort.

Save 20%+ with a team

Private cohort

Run a cohort for your org

A dedicated cohort with a custom schedule and curriculum, tailored to your team.

Book a private cohort

$600

USD

Nov 2
·

12–4:30pm EST

Enroll