
Free Lesson
π§βπ« Teach a Model: Post-Training, RLHF, and Reasoning
Part of Life of a Model: Build, Teach, Tune, Ship
60 min
Nov 5, 2026 2:00 PM
By continuing, you agree to Maven's Terms and Privacy Policy.
What you'll learn
Teach a base model to follow instructions with SFT
See how example conversations turn a text predictor into a model that answers questions and follows directions.
Compare RLHF and DPO for learning human preferences
Learn how human preferences become a training signal, and when each method is the right fit.
See how reasoning RL and red teaming shape behavior
Learn how verifiable rewards build step-by-step thinking and how red teaming finds failures before users do.
Why this topic matters
A base model can complete text, but it can't reliably answer a question, refuse a harmful request, or work through a hard problem. Post-training is where those behaviors are shaped, and it's why ChatGPT, Claude, and Gemini feel so different from raw models. It's also where quality and safety trade-offs get decided, which makes it essential knowledge for anyone judging or choosing a model.





