Free Lesson
Build Multi-Agent Systems You Can Audit
Part of The AI Evaluation Handbook
30 min
Jun 24, 2026 11:00 AM
Virtual (Zoom)
In this video
What you'll learn
What makes a forecast scoreable, not just persuasive
Five stages — question, evidence, independent runs, aggregation, scoring — that decide whether your number holds up.
Audit while the agent works, not after
Replayable evidence, independent runs, explicit aggregation — captured as the trace forms, not bolted on at the end.
The result most agent demos hide
Brier, log score, calibration locked before resolution — including when agent + consensus beats either alone.
Why this topic matters
Most AI forecasts collapse under one question: "Why should I trust this number?" Auditability isn't bolted on at the end — it's decided five stages earlier: the question, the evidence, the independent runs, the aggregation rule, and the scoring plan locked before resolution. Build the process; the audit takes care of itself.
