Founder | Ex-Google | Prof UCLA & UMN
Data Science Leader at Google


Most teams test agents the same way: run them a few times, look at the outputs, and if everything seems reasonable, ship.
“Looks good” is not an evaluation strategy.
The problems show up later. The agent loops, calls the wrong tool, takes too many steps, or gets the right answer through a completely broken path.
In this hands-on lightning lesson, we’ll build a practical approach to evaluating agentic systems. Instead of only checking the final answer, we’ll look at what the agent actually did along the way.
What we'll coverBuilding golden task sets for real agent behavior
Using LLM-as-a-Judge with meaningful rubrics
Evaluating agent trajectories, not just final answers
Testing tool selection and tool calls
Catching loops, context loss, and unsafe actions
Running evals against a real agent harness
You'll leave with an eval set, a working judge, and a repeatable framework you can apply to your own agents.
PrerequisitesNo coding required. We'll provide the example agent and walk through everything together. A basic understanding of LLMs and agents is helpful, but this session is designed for engineers, PMs, technical leaders, and AI builders.
Master AI strategy, agents, and execution, transform from passive observer to confident leader driving high impact AI decisions in your org.
The gap between manual spot-checks and production-scale reliability
The taxonomy of eval types: assertions, LLM-as-judge, human review
Select one high-impact use case from your own team
Understanding failure cases and writing test cases against a shared example app
Covering edge cases and adversarial inputs, not just typical usage
Why 30 diverse cases beat 300 similar ones
When a code-based assertion is enough vs. when you need a model-graded judge
Designing a judge prompt with a clear rubric and forced structured verdicts, then calibrating it against human labels
Catching judge bias: position, verbosity, self-preference
Build a ground truth dataset for your use case
Apply the 8/10 threshold to decide readiness or rejection
Create evaluation criteria to hold systems and vendors accountable
Judging the trajectory: tool calls, reasoning steps, order of operations and not just the final answer
Tool-use correctness, error handling, and termination behavior
Designing a lightweight harness/sandboxed environment for agent evals
Avoid common adoption failures (no owner, no fallback, no feedback)
Communicate AI initiatives effectively to stakeholders
Present a complete, actionable agent blueprint your team can execute
80% of enterprise AI projects fail to reach production—mainly due to wrong use cases, no evaluation, and no adoption plan. Learn a live scoring framework to pick and justify one high-impact use case.
Break down a production AI agent—inputs, tools, memory, outputs—no code. Learn when to use multi-agent pipelines, see a live demo, and map your own workflow on a paper like an org chart.
This hour covers what others skip: build a 10-row ground truth table to evaluate your agent objectively. Learn the 8/10 threshold, when to cut a use case, and how to hold vendors accountable.
Deployment isn’t technical—it’s organizational. Build a 30-day rollout plan: ownership, pilot users, early risks, and week-one metrics. Learn the 3 things that kill adoption and how to avoid them.

Founder Traversaal.ai | Adjunct Professor | 15+ years | Google | Stanford | UCLA

13 Years+ in building Large Scale Data Science Models at Google
Enterprise PMs and product leads who’ve been handed the AI mandate but don’t know where to start. You’ve seen demos. You need a process.
Non-technical operators and team leads who know AI can save their team 10 hours a week but can’t figure out which use case to start with.
Executives and directors who are tired of hearing “it’s almost ready”. You want to evaluate and know whether it’s actually production-grade.
Live sessions
Learn directly from Hamza Farooq & Manisha Arora in a real-time, interactive format.
Live 4-hour build session
Work alongside Hamza and a cohort of enterprise peers. Real exercises, real feedback, real use cases — not slides.
Lifetime access to recordings
Go back to any part of the workshop whenever you onboard a new team member or start a new agent project.
Complete Agent Blueprint Template
The exact one-page format Traversaal uses to scope new agent projects. Covers use case scoring, workflow diagram, eval table, rollout plan, and success metrics. Ready to present to your leadership team.
Ground Truth Eval Rubric
A 10-row eval table template with scoring guide. Use it before you build, after you build, and every time you bring in a vendor. The only objective way to know if your agent is production-ready.
30-Day Rollout Playbook
A week-by-week checklist for taking an agent from internal dogfood to team-wide deployment. Includes the owner assignment matrix, feedback loop setup, and the 3 early warning signs that your rollout is stalling.
Vendor Accountability Checklist
Know exactly what to ask any AI vendor before signing. Based on what Traversaal has seen fail in production. Most enterprise buyers never ask these questions until it’s too late.
Community of peers
Connect with enterprise PMs, operators, and leads facing the same challenges. Share what’s working, compare evals, and get feedback on your blueprint before you pitch it internally.
Certificate of completion
Share your new skills with your employer or on LinkedIn.
Maven Guarantee
Your purchase is backed by the Maven Guarantee.
Maven for Teams
Reimbursement
Get your company to pay
Everything L&D needs: email template, receipts, and certificate of completion.
Get reimbursedTeam discount
Learn with your teammates
Save 20%+ when 2 or more teammates enroll in the same cohort.
Save 20%+ with a teamPrivate cohort
Run a cohort for your org
A dedicated cohort with a custom schedule and curriculum, tailored to your team.
Book a private cohort$500
USD
11:30am–3:30pm EDT