The Future of LLM Serving

Hosted by Abi Aryan

Mon, Aug 24, 2026

3:30 PM UTC (45 minutes)

Virtual (Zoom)

Free to join

74 students

Invite your network

Go deeper with a course

AI Inference Engineering & Systems Design
Abi Aryan
View syllabus

What you'll learn

What actually makes inference fast?

Understand where cost and latency comes from and what really affects how quickly a model can respond.

Why does inference get so expensive?

See what drives serving costs as models get larger, traffic grows, and users expect faster responses.

What changes at production scale?

Explore why serving one request is easy, but serving thousands efficiently is a completely different problem.

Why this topic matters

LLM inference is where AI becomes a real product. It determines not only how fast, and scalable an AI experience can be but most importantly can you afford it. Without understanding this, you can sit with the sexiest RAG or Agentic applications, but they are just demos and will fail the moment real traffic hits.

You'll learn from

Abi Aryan

Computer Scientist and ML Engineer

Abi Aryan is the ex-founder of Abide AI and a machine learning engineer with over a decade of experience building production-level ML systems.

A mathematician by training, she previously was a visiting research scholar at UCLA, supervised by the 2012 Turing Award winner, Dr. Judea Pearl, where she focused on developing intelligent agents.

Abi has authored research papers in AutoML, multi-agent systems, and large language models, and actively reviews for leading research conferences and workshops, including NeurIPS, ACL, EMNLP, and AABI

She is currently teaching and advancing doctoral research at the intersection of several key areas:

• Distributed systems

• GPU engineering for large-scale AI systems

• SLO-Aware Inference Optimization

• Low-cost chip design

• AI Policy

Author

O'Reilly Media
See all products from Abi (@goabiaryan)

Sign up to join this lesson

By continuing, you agree to Maven's Terms and Privacy Policy.