Build Your First AI Eval

Jules Boiteux

Ex-PM turned AI-Native Product Builder

The course for AI-native PMs who want to build their first AI eval from scratch

🚨 Save 20% with this link

----

💰 The most hands-on AI evals course you'll find under $500

----

AI evals are the hottest skill an AI-native PM can have: more and more features run on an LLM, and their accuracy keeps competitors at bay.

The problem is where to start. Evals feel like an engineer's job, the information is scattered, the courses long and expensive.

Concrete practice, not slides. You run every eval on a live example: an agent sorting the real feature requests of Excalidraw, an open-source whiteboard app.

In 3 sessions of 1.5 hours, you go from zero to your own eval tool: a golden set, a scored system prompt, an LLM as a judge, then a tool you vibe code to run your AI evals.

No coding, no Claude Code knowledge required: the method works with any tool (Codex, Cursor...). I teach with Claude Code, the easiest way to apply them. We set up together in session 1.

----

🛠️ 60% hands-on practice, 40% theory

📚 Lifetime access to recordings and future cohorts

✍️ Written review of your golden set and judge

🧪 A concrete example to replicate for any AI eval you run

🗂️ A GitHub boilerplate for your own evals

👥 A private Slack channel, even after the course

What you’ll learn

By the end, you can apply a robust and replicable evaluation method to any AI feature, as the PM who manages it or the Builder who ships it

  • What makes a good golden set of data

  • How to come up with the right criteria for your golden set

  • A real example: automatically categorizing Excalidraw's feature requests

  • The minimal tool stack for your first AI evals, low budget and few dependencies

  • A ready-made script that connects to hundreds of models via Vercel AI Gateway

  • How to visualize your results

  • How to derive the system prompt for your first run

  • How to analyze the results from your first run

  • How to leverage manual analysis to improve your system prompt

  • When you need an LLM as a judge

  • How to write a judge and check it against your own grades

  • Why you freeze the judge before improving the prompt

  • What an eval tool needs: grading, runs and history

  • How to plan it with Claude Code before any code is written

  • How to share your results with your dev team

  • How to reuse the method on your own AI feature

  • How to keep your evals honest over time

  • How to explain your results to your team and stakeholders

Learn directly from Jules

Jules Boiteux

Jules Boiteux

Former Senior Product Manager, now founder of Vibe Coding Academy & AI educator

Trained Teams at
Instacart
Amazon
PayPal
Optimizely
GoDaddy
See all products from Jules

Who this course is for

  • Product Managers, Product Owners and AI-native Product Builders who ship AI features and want proof that a change helped

  • Software engineers building LLM features who want a score next to every prompt version instead of a gut feeling

  • Leaders who steer product and tech teams on AI features, who want to talk evals with them, challenge on evidence and set priorities.

Prerequisites

  • A subscription to Claude, OpenAI or Cursor

    Taught with Claude Code, so Claude Pro is smoothest; Codex or Cursor Pro work too. Plus a laptop you can install on.

  • Basic AI literacy (no coding required)

    You have used Claude or ChatGPT and know what a prompt and an API are. No coding: session 3 is vibe coding, you direct.

  • No AI feature of your own needed

    We run every eval on Excalidraw's public issues, taught as a method: each step maps to your own AI feature, when you are ready.

What's included

Jules Boiteux

Live sessions

Learn directly from Jules Boiteux in a real-time, interactive format.

Personal written review

Jules reviews your golden set, your prompt and your judge after every session, with feedback specific to your own AI feature.

Lifetime access to course material

Every deck, plus the resources doc with ready-to-use prompts, step-by-step process instructions and the criteria templates.

Community of peers

Stay accountable and share insights with like-minded professionals.

Certificate of completion

Share your new skills with your employer or on LinkedIn.

Private cohort Slack channel

Ask questions, compare scores and unblock between sessions.

Lifetime access to recordings

Every session is recorded and stays yours for life.

A GitHub repo with the structure of a project you can learn from

A ready-to-clone repo with the project structure, the scripts to run any model through Vercel AI Gateway, the criteria and prompt templates, and the judge scaffold. Swap in your own data and you have your first eval running the same day.

Maven Guarantee

Your purchase is backed by the Maven Guarantee.

Course syllabus

5 live sessions • 16 lessons • 3 projects

Week 1

Nov 10—Nov 15

    Nov

    10

    Session 1 - Run your first AI Eval

    Tue 11/105:00 PM—6:30 PM (UTC)

    Your First AI Eval: Categorizing Excalidraw's Feature Requests

    7 items

    Nov

    13

    Office Hour 1: Q&A on Your First Eval

    Fri 11/134:30 PM—5:00 PM (UTC)

Week 2

Nov 16—Nov 22

    Nov

    17

    Session 2 - Your Second AI Eval, with an LLM as a Judge

    Tue 11/175:00 PM—6:30 PM (UTC)

    Grading Free Text: an Automatic Reply to Excalidraw's Issue Authors

    7 items

    Nov

    20

    Office Hour 2: Q&A on Your LLM Judge

    Fri 11/204:30 PM—5:00 PM (UTC)

Schedule

Live sessions

1-2 hrs / week

Each live session last 1h30

    • Tue, Nov 10

      5:00 PM—6:30 PM (UTC)

    • Fri, Nov 13

      4:30 PM—5:00 PM (UTC)

    • Tue, Nov 17

      5:00 PM—6:30 PM (UTC)

Projects

1-2 hrs / week

After each session, an assignment applies the method to your own AI feature. Optional, but this is where it sticks. Jules reviews every submission and sends written, personalized feedback.

Session replays

1 hr / week

Every session is recorded. Rewatch the replay when you want to review a concept from class or catch up on a session you missed. Optional: about 30 minutes a week, depending on how much you want to revisit.

Frequently asked questions

Maven for Teams

Reimbursement

Get your company to pay

Everything L&D needs: email template, receipts, and certificate of completion.

Get reimbursed

Team discount

Learn with your teammates

Save 20%+ when 2 or more teammates enroll in the same cohort.

Save 20%+ with a team

Private cohort

Run a cohort for your org

A dedicated cohort with a custom schedule and curriculum, tailored to your team.

Book a private cohort

$499

USD

Nov 10—Nov 23
Enroll