Priyanka Vergadia
Free Lesson

πŸ§‘β€πŸ« Teach a Model: Post-Training, RLHF, and Reasoning

Part of Life of a Model: Build, Teach, Tune, Ship

60 min
Nov 5, 2026 2:00 PM

By continuing, you agree to Maven's Terms and Privacy Policy.

What you'll learn

Teach a base model to follow instructions with SFT

See how example conversations turn a text predictor into a model that answers questions and follows directions.

Compare RLHF and DPO for learning human preferences

Learn how human preferences become a training signal, and when each method is the right fit.

See how reasoning RL and red teaming shape behavior

Learn how verifiable rewards build step-by-step thinking and how red teaming finds failures before users do.

Why this topic matters

A base model can complete text, but it can't reliably answer a question, refuse a harmful request, or work through a hard problem. Post-training is where those behaviors are shaped, and it's why ChatGPT, Claude, and Gemini feel so different from raw models. It's also where quality and safety trade-offs get decided, which makes it essential knowledge for anyone judging or choosing a model.

You'll learn from

Priyanka Vergadia

Priyanka Vergadia

Visual Product Storyteller | Built & Shipped AI at Google & Microsoft

Previously at & Trusted By:
Google
Microsoft
The Wharton School
Intel
TED
See all products from Priyanka
Get free access