
Build Better AI Agents: Architecture, Harnesses, and Evals
Your agent looks good in a demo. Then it picks the wrong tool, loses context, or claims it finished work it never did. Across two free lessons, learn to avoid five common agent-building mistakes, design a harness around your task, and choose evals that catch failures and verify improvements. Learn what changes across coding, research, conversational, and computer-use agents.
By continuing, you agree to Maven's Terms and Privacy Policy.
Thu Sep 17·11:00 PM UTC
Five Mistakes Everyone Makes When Building AI Agents
You add memory because coding agents have it, give the model control over steps your code could handle, and keep every workaround when you upgrade the model. Then you change the prompt, tools, and architecture together and wonder what helped. I'll unpack five mistakes in agent architecture, harness engineering, and evaluation, showing how to build less, inspect failures, and verify improvements.
You'll learn from

Hugo Bowne-Anderson
AI & data engineer, consultant, educator of 6+ million students (ex-Yale)
Mon Sep 21·10:00 PM UTC
AI Agent Evals: Test What Matters for Your Agent
Your research agent cites sources that don't support its claims. Your support agent says it booked an appointment, but nothing changed in the system. A generic quality score can hide both failures. I'll cover the foundations of agent evals, show how success criteria change across agent types, and explain how to use the results to improve your harness without breaking what already works.
You'll learn from

Hugo Bowne-Anderson
AI and data scientist, consultant, educator of 6+ million students (ex-Yale)