Building AI Evals in Practice: Agent Trajectories and Harness Design

Hamza Farooq

Founder | Ex-Google | Prof UCLA & UMN

Manisha Arora

Data Science Leader at Google

Your Agent Works. But Can You Trust It?

Most teams test agents the same way: run them a few times, look at the outputs, and if everything seems reasonable, ship.

“Looks good” is not an evaluation strategy.

The problems show up later. The agent loops, calls the wrong tool, takes too many steps, or gets the right answer through a completely broken path.

In this hands-on lightning lesson, we’ll build a practical approach to evaluating agentic systems. Instead of only checking the final answer, we’ll look at what the agent actually did along the way.

What we'll cover
  • Building golden task sets for real agent behavior

  • Using LLM-as-a-Judge with meaningful rubrics

  • Evaluating agent trajectories, not just final answers

  • Testing tool selection and tool calls

  • Catching loops, context loss, and unsafe actions

  • Running evals against a real agent harness

You'll leave with an eval set, a working judge, and a repeatable framework you can apply to your own agents.

Prerequisites

No coding required. We'll provide the example agent and walk through everything together. A basic understanding of LLMs and agents is helpful, but this session is designed for engineers, PMs, technical leaders, and AI builders.

What you’ll learn

Master AI strategy, agents, and execution, transform from passive observer to confident leader driving high impact AI decisions in your org.

  • The gap between manual spot-checks and production-scale reliability

  • The taxonomy of eval types: assertions, LLM-as-judge, human review

  • Select one high-impact use case from your own team

  • Understanding failure cases and writing test cases against a shared example app

  • Covering edge cases and adversarial inputs, not just typical usage

  • Why 30 diverse cases beat 300 similar ones

  • When a code-based assertion is enough vs. when you need a model-graded judge

  • Designing a judge prompt with a clear rubric and forced structured verdicts, then calibrating it against human labels

  • Catching judge bias: position, verbosity, self-preference

  • Build a ground truth dataset for your use case

  • Apply the 8/10 threshold to decide readiness or rejection

  • Create evaluation criteria to hold systems and vendors accountable

  • Judging the trajectory: tool calls, reasoning steps, order of operations and not just the final answer

  • Tool-use correctness, error handling, and termination behavior

  • Designing a lightweight harness/sandboxed environment for agent evals

  • Avoid common adoption failures (no owner, no fallback, no feedback)

  • Communicate AI initiatives effectively to stakeholders

  • Present a complete, actionable agent blueprint your team can execute

Workshop agenda

  • The Pilot Graveyard: Why Enterprise AI Fails Before It Ships

    80% of enterprise AI projects fail to reach production—mainly due to wrong use cases, no evaluation, and no adoption plan. Learn a live scoring framework to pick and justify one high-impact use case.

  • Design Your Agent: From Workflow to Blueprint

    Break down a production AI agent—inputs, tools, memory, outputs—no code. Learn when to use multi-agent pipelines, see a live demo, and map your own workflow on a paper like an org chart.

  • Evals: How to Know If It’s Actually Working

    This hour covers what others skip: build a 10-row ground truth table to evaluate your agent objectively. Learn the 8/10 threshold, when to cut a use case, and how to hold vendors accountable.

  • Roll It Out: Getting Your Team to Actually Use It

    Deployment isn’t technical—it’s organizational. Build a 30-day rollout plan: ownership, pilot users, early risks, and week-one metrics. Learn the 3 things that kill adoption and how to avoid them.

Learn directly from Hamza & Manisha

Hamza Farooq

Hamza Farooq

Founder Traversaal.ai | Adjunct Professor | 15+ years | Google | Stanford | UCLA

Manisha Arora

Manisha Arora

13 Years+ in building Large Scale Data Science Models at Google

Google
See all products from Hamza

Who this workshop is for

  • Enterprise PMs and product leads who’ve been handed the AI mandate but don’t know where to start. You’ve seen demos. You need a process.

  • Non-technical operators and team leads who know AI can save their team 10 hours a week but can’t figure out which use case to start with.

  • Executives and directors who are tired of hearing “it’s almost ready”. You want to evaluate and know whether it’s actually production-grade.

What's included

Live sessions

Learn directly from Hamza Farooq & Manisha Arora in a real-time, interactive format.

Live 4-hour build session

Work alongside Hamza and a cohort of enterprise peers. Real exercises, real feedback, real use cases — not slides.

Lifetime access to recordings

Go back to any part of the workshop whenever you onboard a new team member or start a new agent project.

Complete Agent Blueprint Template

The exact one-page format Traversaal uses to scope new agent projects. Covers use case scoring, workflow diagram, eval table, rollout plan, and success metrics. Ready to present to your leadership team.

Ground Truth Eval Rubric

A 10-row eval table template with scoring guide. Use it before you build, after you build, and every time you bring in a vendor. The only objective way to know if your agent is production-ready.

30-Day Rollout Playbook

A week-by-week checklist for taking an agent from internal dogfood to team-wide deployment. Includes the owner assignment matrix, feedback loop setup, and the 3 early warning signs that your rollout is stalling.

Vendor Accountability Checklist

Know exactly what to ask any AI vendor before signing. Based on what Traversaal has seen fail in production. Most enterprise buyers never ask these questions until it’s too late.

Community of peers

Connect with enterprise PMs, operators, and leads facing the same challenges. Share what’s working, compare evals, and get feedback on your blueprint before you pitch it internally.

Certificate of completion

Share your new skills with your employer or on LinkedIn.

Maven Guarantee

Your purchase is backed by the Maven Guarantee.

Free resources

Frequently asked questions

Maven for Teams

Reimbursement

Get your company to pay

Everything L&D needs: email template, receipts, and certificate of completion.

Get reimbursed

Team discount

Learn with your teammates

Save 20%+ when 2 or more teammates enroll in the same cohort.

Save 20%+ with a team

Private cohort

Run a cohort for your org

A dedicated cohort with a custom schedule and curriculum, tailored to your team.

Book a private cohort

$500

USD

·
Oct 9
·

11:30am–3:30pm EDT

Enroll