Free Lesson
Make AI Agents Fail Safely
45 min
Aug 11, 2026 11:00 AM
Virtual (Zoom)
In this video
What you'll learn
Classify failures before selecting a response
Learn to distinguish among:
transient failures;
permanent failures;
missing or invalid data;
model-quality failure
Add timeouts and bounded retries
See why unlimited retries can make an incident slower, more expensive and harder to diagnose
Build a practical fallback hierarchy
The safest fallback is not always another model. In some cases, the correct production behaviour is to avoid acting.
Preserve workflow state
Understand how explicit states & checkpoints allow an agent to continue from the last successful step vs restarting
Why this topic matters
A production AI agent is connected to models, APIs, databases, enterprise tools and human approval processes. Each dependency can fail independently.
Without reliability engineering, one can get:
incorrect prices or decisions;
duplicate transactions;
compliance exposure;
poor customer experiences.
A production-ready agent must therefore be judged by both its success and its failure behavior
You'll learn from

Dr Ankur Narang
Dr. Ankur Narang brings 30+ yrs exp in AI, Tech across MNCs & many verticals

Kush Khurana
AI & ML leader; Venture Partner, DeepCoreX; Ashoka faculty
IBM; Oracle; Meta; Apparel Group; Hike