Avery Yip
Lucy Chen
Colin Matthews
Free Lesson

The Hillclimb: How to Measure and Improve Your AI Agents

60 min
Sep 16, 2026 3:00 PM

By continuing, you agree to Maven's Terms and Privacy Policy.

What you'll learn

Demystify evals

What evals, benchmarks, and training datasets are, and how they fit together for your product.

Design structured evals

Taxonomize workflows that matter, write verifiable rubric criteria, and build simulations that check what your agent did

Complete the learning loop

Turn failure modes into training data your agent hillclimbs on, so evals become compounding IP instead of a report card.

Why this topic matters

Most teams check their agents on vibes: a few test prompts and sparse top-line metrics. Structured Evals systematically identify where your agent fails, with simulations that mimic real-world scenarios, and turn those failures into the training data that teaches the agent to do the work better. That's how you build AI you can trust. In this session, we'll walk through a real case study: an agent that automates pay disputes.

You'll learn from

Avery Yip

Avery Yip

Director of Engineering, Handshake AI

Lucy Chen

Lucy Chen

Head of Learning and Training, Handshake AI

Colin Matthews

Colin Matthews

Head of Education, Lenny's Newsletter

See all products from Colin
Get free access