Free Lesson

The Future of LLM Serving

45 min
Aug 27, 2026 11:30 AM
Virtual (Zoom)

In this video

What you'll learn

What actually makes inference fast?

Understand where cost and latency comes from and what really affects how quickly a model can respond.

Why does inference get so expensive?

See what drives serving costs as models get larger, traffic grows, and users expect faster responses.

What changes at production scale?

Explore why serving one request is easy, but serving thousands efficiently is a completely different problem.

Why this topic matters

LLM inference is where AI becomes a real product. It determines not only how fast, and scalable an AI experience can be but most importantly can you afford it. Without understanding this, you can sit with the sexiest RAG or Agentic applications, but they are just demos and will fail the moment real traffic hits.

You'll learn from

Abi Aryan

Abi Aryan

Computer Scientist and ML Engineer

Author

O'Reilly Media
See all products from Abi (@goabiaryan)