Free Lesson

Build Multi-Agent Systems You Can Audit

Part of The AI Evaluation Handbook

30 min
Jun 24, 2026 11:00 AM
Virtual (Zoom)

In this video

What you'll learn

What makes a forecast scoreable, not just persuasive

Five stages — question, evidence, independent runs, aggregation, scoring — that decide whether your number holds up.

Audit while the agent works, not after

Replayable evidence, independent runs, explicit aggregation — captured as the trace forms, not bolted on at the end.

The result most agent demos hide

Brier, log score, calibration locked before resolution — including when agent + consensus beats either alone.

Why this topic matters

Most AI forecasts collapse under one question: "Why should I trust this number?" Auditability isn't bolted on at the end — it's decided five stages earlier: the question, the evidence, the independent runs, the aggregation rule, and the scoring plan locked before resolution. Build the process; the audit takes care of itself.

You'll learn from

Stefan Jansen

Stefan Jansen

Author, ML for Trading · Founder, Applied AI · Investing since 2013

See all products from Stefan