Free Lesson
Evaluating LLMs Beyond Accuracy: Pragmatic Reasoning
45 min
Jul 31, 2026 12:00 PM
Virtual (Zoom)
In this video
What you'll learn
Spot the accuracy illusion
Learn why high benchmark scores can mask genuine reasoning failures in LLMs
Understand pragmatic reasoning
See how context, implication, and inference reveal what models reasoning actually process
Identify shortcut-taking patterns
Recognize when a model lands on the right answer via the wrong path
Apply sharper eval criteria
Use research-backed methods to stress-test LLM reasoning in your own work
Why this topic matters
LLMs are being deployed in high-stakes contexts based on benchmark scores that mask real reasoning gaps. Understanding how models exploit statistical shortcuts rather than genuine pragmatic inference is critical for anyone building or evaluating AI systems. This talk gives practitioners a sharper lens for assessing when to trust LLM outputs.
You'll learn from

Amir Feizpour
Founder @ Aggregate Intellect

Tara Azin
PhD Candidate @ Carleton University (Language & Logic Lab)
