Free Lesson

Evaluating LLMs Beyond Accuracy: Pragmatic Reasoning

45 min
Jul 31, 2026 12:00 PM
Virtual (Zoom)

In this video

What you'll learn

Spot the accuracy illusion

Learn why high benchmark scores can mask genuine reasoning failures in LLMs

Understand pragmatic reasoning

See how context, implication, and inference reveal what models reasoning actually process

Identify shortcut-taking patterns

Recognize when a model lands on the right answer via the wrong path

Apply sharper eval criteria

Use research-backed methods to stress-test LLM reasoning in your own work

Why this topic matters

LLMs are being deployed in high-stakes contexts based on benchmark scores that mask real reasoning gaps. Understanding how models exploit statistical shortcuts rather than genuine pragmatic inference is critical for anyone building or evaluating AI systems. This talk gives practitioners a sharper lens for assessing when to trust LLM outputs.

You'll learn from

Amir Feizpour

Amir Feizpour

Founder @ Aggregate Intellect

Tara Azin

Tara Azin

PhD Candidate @ Carleton University (Language & Logic Lab)

See all products from aggregate