Free Lesson
Scaling Late Interaction to Billions of Documents
Part of AI Product Engineering
45 min
Jul 9, 2026 2:00 PM
Virtual (Zoom)
In this video
What you'll learn
Why single-vector retrieval loses detail
Understand what gets collapsed when each document is represented as one vector, and why that hurts long-tail queries
What late interaction changes
See how token- and patch-level representations preserve more semantic detail than single-vector retrieval.
Why late interaction is expensive
Learn where the 10-100x storage and compute overhead comes from.
How we scaled it to billions of document
Learn the production methods used to keep late interaction retrieval fast at billion-document scale.
How production constraints shape retrieval
See how sub-100ms p99 latency, filtering, and online index updates affect the architecture.
Why this topic matters
Single-vector retrieval is cheap, but it throws away detail that matters for hard queries. Late interaction keeps
more of that detail, but the production cost is large. This lesson shows how we scaled it to billions of documents
while keeping latency, filtering, and index updates practical.
You'll learn from

Marek Galovic
CEO, Co-Founder @TopK. ex-Pinecone, ex-Shopify

Hamel Husain
ML Engineer with 20+ years of experience
Previously at