Free Lesson

Scaling Late Interaction to Billions of Documents

Part of AI Product Engineering

45 min
Jul 9, 2026 2:00 PM
Virtual (Zoom)

In this video

What you'll learn

Why single-vector retrieval loses detail

Understand what gets collapsed when each document is represented as one vector, and why that hurts long-tail queries

What late interaction changes

See how token- and patch-level representations preserve more semantic detail than single-vector retrieval.

Why late interaction is expensive

Learn where the 10-100x storage and compute overhead comes from.

How we scaled it to billions of document

Learn the production methods used to keep late interaction retrieval fast at billion-document scale.

How production constraints shape retrieval

See how sub-100ms p99 latency, filtering, and online index updates affect the architecture.

Why this topic matters

Single-vector retrieval is cheap, but it throws away detail that matters for hard queries. Late interaction keeps more of that detail, but the production cost is large. This lesson shows how we scaled it to billions of documents while keeping latency, filtering, and index updates practical.

You'll learn from

Marek Galovic

Marek Galovic

CEO, Co-Founder @TopK. ex-Pinecone, ex-Shopify

Hamel Husain

Hamel Husain

ML Engineer with 20+ years of experience

Previously at

Pinecone
GitHub
Shopify.com
Airbnb
See all products from Hamel Husain & Shreya Shankar