

Free Lesson
Optimizing Open Models for Production
Part of Global AI Builder Series: by The Gen Academy
60 min
Sep 1, 2026 2:00 PM
What you'll learn
Understand Modern Inference Optimization
Learn the techniques that improve latency, throughput, memory efficiency, and cost when serving large open models.
Optimize Frontier Open Models at Scale
Explore caching, quantization, optimized kernels, disaggregated serving, & speculative decoding for production workloads
Learn from the Kimi K3 Production Stack
See how recent inference optimizations helped bring the 2.8T-parameter Kimi K3 model into production efficiently.
Why this topic matters
Open models are getting dramatically larger and more capable, which makes serving them efficiently a systems problem of its own.
This session explores modern inference techniques used to make state-of-the-art open models practical in production, covering caching, kernels, quantization, speculative decoding, and disaggregated inference, with lessons from the recent Kimi K3 launch and its production-scale vLLM integration.






