AI Evals and Reliability for Product Leaders

Dr. Aki Wijesundara

PhD in ML | Google AI Accelerator Alum

Manu Jayawardana

Exited AI Founder | Founder, TAI Labs

Early Bird Discount ends in 24 hours, use the code EVALS60 for 60% OFF

You already ship an AI surface. You also know the honest answer to "is it actually working?" is a dashboard nobody trusts and a demo that behaves on stage.

Meanwhile the roles you want, at frontier labs and AI-first companies, screen for one thing: have you personally owned an AI quality decision. Not shipped near one. Owned it.

In 4 weeks you'll build the evidence.

You'll walk away with four portfolio artifacts:

  • Build a failure taxonomy from 100+ traces

  • Calibrate an LLM judge and report a corrected number with a confidence interval

  • Red-team a live agent and write the launch decision memo

  • Present and defend a full quality strategy to the cohort

Our students come from MAANG companies, and they lead AI products at scale.

Want to see how we teach first? Watch our free recorded mini-series, Think Like an AI Engineer.

A limited number of need-based scholarships are available for applicants facing real financial hardship. See if you qualify.

What you’ll learn

You’ll create an AI Evals Launch Pack for a real feature you’re shipping

  • Read 100 supplied traces and code failures the way qualitative researchers do, open coding then axial coding

  • Turn raw failure notes into named, countable categories your engineers can act on

  • Write the prioritisation memo that decides where engineering time goes

  • Design pass/fail criteria that survive contact with a real labelling session

  • Measure your judge against your own labels, then correct the estimate and put a confidence interval on it

  • Explain to a skeptical VP why a 30-item eval set tells you almost nothing

  • Run structured red-teaming against a live broken agent, not a vibes exercise

  • Work prompt injection, goal hijacking, tool poisoning, and data exfiltration

  • Produce a findings report with severity, blast radius, and residual risk you would accept

  • Separate what belongs at runtime as a guardrail from what belongs offline as an eval

  • Set logging, sampling, and regression gates that catch drift after a model upgrade

  • Treat cost and latency as quality dimensions, not someone else's problem

  • Write the quality section of a PRD that engineering actually respects

  • Run an incident review for an AI failure and set a bar that survives a launch deadline

  • Decide who owns evals and how to get the org to fund it

  • Position a big-tech PM background for frontier lab and AI-first product roles

  • Work the evals interview question bank live, with strong and weak answers modelled

  • Defend your capstone in a mock case-study round with written feedback

Learn directly from Aki & Manu

Dr. Aki Wijesundara

Dr. Aki Wijesundara

AI Founder | Educator | Google AI Accelerator Alum

Google
Meta
Amazon Web Services
OpenAI
NVIDIA
Manu Jayawardana

Manu Jayawardana

AI Advisor | Founder, TAI Labs

Previous Students from
OpenAI
Boston Consulting Group (BCG)
NVIDIA
Google
McKinsey & Company
See all products from TAI Labs

Who this course is for

  • The Senior or Lead PM. Already shipping an AI surface. Tired of guessing whether it works and answering for it anyway.

  • The Big-Tech PM Going Frontier. At Uber, Meta, or Google. Wants safety, evals, or trust roles and needs proof, not vocabulary.

  • The AI Product Leader. Owns a team shipping AI. Needs a quality function and a defensible bar, not another dashboard.

Prerequisites

  • You currently work on or near an AI product

    The work assumes you have shipped or are shipping an AI feature. Bring your own context, or use the supplied corpus.

  • Comfort reading data, no coding required

    You will edit prompts and read notebook output. All infrastructure is scaffolded and hosted for you.

  • Willingness to be graded on your judgement

    Every assignment is a decision you defend in writing. The feedback on your reasoning is the point.

What's included

Live sessions

Learn directly from Dr. Aki Wijesundara & Manu Jayawardana in a real-time, interactive format.

4 live workshop sessions

Two hours each, hands-on. We read traces, break agents, and calibrate judges together in real time.

Three graded portfolio artifacts

A failure taxonomy, a calibrated judge, and a red-team report with a launch decision. Each one is showable in an interview.

Written feedback on your reasoning

Small cohort so every assignment gets human review. You are graded on judgement quality, against a rubric published on day one.

Capstone defence and mock case round

Present a full quality strategy to the cohort and defend it live, the way a real case-study interview runs.

Interview question bank and portfolio review

The evals questions actually asked at frontier labs, plus written feedback on how your three artifacts land as interview material.

Guest practitioner session

A session with someone doing evals or safety work at a lab today, including what they screen for when hiring.

Office hours and lifetime access

Weekly Q&A plus permanent access to recordings, rubrics, templates, and the trace corpus.

Maven Guarantee

Your purchase is backed by the Maven Guarantee.

Course syllabus

4 live sessions • 4 lessons • 4 projects

Week 1

Aug 31—Sep 6

    Aug

    31

    Week 1: See the failures

    Mon 8/313:00 PM—5:00 PM (UTC)

    Week 1 : See the Failures

    2 items

Week 2

Sep 7—Sep 13

    Sep

    7

    Week 2: Measure the failures honestly

    Mon 9/73:00 PM—5:00 PM (UTC)

    Week 2 : Measure the Failures Honestly

    2 items

Free resource

Build an App with AI in 30 Minutes cover image

Build an App with AI in 30 Minutes

Idea → Working Code (Cursor)

Watch Cursor agents turn plain-English prompts into real code and refactor it on the fly.

Pixel-Perfect UI in Seconds (Lovable)

Generate a slick, production-ready frontend without touching Figma.

Ship Before Your Coffee Cools

Glue the two tools together, deploy, and walk away with a live link—all inside half an hour.

Schedule

Live sessions

2 hrs

One two-hour workshop each week. We work on real material together, reading traces, demonstrating judge bias live, and running attacks against a broken agent. Every session ends with the assignment framed and the rubric explained, so you know exactly what a strong submission looks like.

    • Mon, Aug 31

      3:00 PM—5:00 PM (UTC)

    • Mon, Sep 7

      3:00 PM—5:00 PM (UTC)

    • Mon, Sep 14

      3:00 PM—5:00 PM (UTC)

    • Mon, Sep 21

      3:00 PM—5:00 PM (UTC)

Projects

4-6 hrs

Four to six hours a week producing one graded decision artifact. Week 1 a failure taxonomy, week 2 a calibrated judge, week 3 a red-team report and launch memo, week 4 a capstone quality strategy you present and defend to the cohort.

Async content

1 hr

Short pre-session primers on the week's concepts, plus the reference pack: rubrics, templates, the interview question bank, and a curated reading list of the papers and practitioner writing that actually matter.

Testimonials

  • The AI training approach is outstanding. Our team learned to build practical AI solutions that we could implement immediately in our educational platform. The hands-on methodology made complex AI concepts accessible to our entire development team.
    Testimonial author image

    Kavi T.

    CEO of Tilli Kids / Stanford PhD
  • Not only are the instructors experts in their field, they're incredibly skilled at breaking down complicated AI concepts so students can grasp them quickly. Anyone interested in building foundational AI knowledge should take this training - it's worth the investment.
    Testimonial author image

    Dr. Elizabeth Creighton

    Founder & Principal at Brazen
  • The instructors help break down AI model development and clearly have plenty of experience to help others learn about complex concepts like infrastructure setup. The practical approach to NLP and LLM applications was exactly what our team needed.
    Testimonial author image

    Alissa Valentine

    NLP & LLM Real World Data Scientist
  • I sent my team through this training for upskilling, and the results have been remarkable. Within weeks, they became much more efficient at building automations and deploying AI agents at work. This program bridges the gap between theory and practice and it’s had a real impact on our productivity.
    Testimonial author image

    Aamir Faaiz

    CEO of Bayseian

Hear It From Our Students

Learning AI Made Simple | Student Feedback on Our AI Engineering Bootcamp | TAI

Our community of learners

A single snapshot of learners across our AI courses and programs.

Student Experience with Our Programs

Student Experience with Our Programs

Student Experience with Our Programs

Student Experience with Our Programs

Student Experience with Our Programs

Frequently asked questions

Maven for Teams

Reimbursement

Get your company to pay

Everything L&D needs: email template, receipts, and certificate of completion.

Get reimbursed

Team discount

Learn with your teammates

Save 20%+ when 2 or more teammates enroll in the same cohort.

Save 20%+ with a team

Private cohort

Run a cohort for your org

A dedicated cohort with a custom schedule and curriculum, tailored to your team.

Book a private cohort

$999

USD

Aug 31Sep 21
Enroll