
Free Lesson
Put Error Bars on Your LLM Metrics
30 min
Sep 23, 2026 2:00 PM
What you'll learn
How to bootstrap a confidence interval on any metric
Resample your results a thousand times in twenty lines of Python. No distribution assumptions required.
How to put error bars on accuracy, cost, and latency
One method covers every number you report. Intervals turn single scores into ranges you can defend.
How to explain score swings between identical runs
Temperature, sampling, and judge variance move scores run to run. See the spread and stop chasing ghosts.
Why this topic matters
You ran the same eval twice and got 84.2, then 81.9. Which number goes in the report? Without error bars, every score is a coin flip dressed as a fact. Teams chase phantom regressions, celebrate phantom wins, and burn weeks on noise. The bootstrap fixes this with twenty lines of Python. Once your metrics carry intervals, your reports survive scrutiny and your decisions stop wobbling.
You'll learn from
Previously at





