Aish & Arvind
Sujee Maniyam
Free Lesson

Optimizing Open Models for Production

Part of Global AI Builder Series: by The Gen Academy

60 min
Sep 1, 2026 2:00 PM

What you'll learn

Understand Modern Inference Optimization

Learn the techniques that improve latency, throughput, memory efficiency, and cost when serving large open models.

Optimize Frontier Open Models at Scale

Explore caching, quantization, optimized kernels, disaggregated serving, & speculative decoding for production workloads

Learn from the Kimi K3 Production Stack

See how recent inference optimizations helped bring the 2.8T-parameter Kimi K3 model into production efficiently.

Why this topic matters

Open models are getting dramatically larger and more capable, which makes serving them efficiently a systems problem of its own. This session explores modern inference techniques used to make state-of-the-art open models practical in production, covering caching, kernels, quantization, speculative decoding, and disaggregated inference, with lessons from the recent Kimi K3 launch and its production-scale vLLM integration.

You'll learn from

Aish & Arvind

Aish & Arvind

AI Engineers | Building & Teaching @ The Gen Academy

Sujee Maniyam

Sujee Maniyam

Developer Relations @ Nebius

See all products from Aishwarya Srinivasan & Arvind Narayan
Get free access