vLLM & SGLang Engineering: End-to-End LLM Serving

Dr. Sreedath Panat

MIT PhD | Vizuara Co-founder

Dr. Rajat Dandekar

Vizuara Co-founder | Purdue PhD

+ Dr. Raj Dandekar

Benchmark, tune and deploy LLMs with vLLM and SGLang

Choosing an inference engine should come from measurements. Learn how vLLM and SGLang serve the same model, where their architectures differ, and which settings improve latency, throughput and GPU cost for your workload.

Across twelve live, hands-on lectures, lead instructor Dr. Sreedath Panat takes you from vLLM V1 and SGLang SRT internals to structured outputs, cache and scheduler tuning, quantization, multi-GPU serving, speculative decoding and production deployment. He is joined by Vizuara co-founders Dr. Raj Dandekar and Dr. Rajat Dandekar. Inference fundamentals are covered in self-paced pre-work.

Build one growing repository: a shared benchmark harness, tuned deployments on both engines, and quantized checkpoints. For your capstone, benchmark the same model on both, deploy your chosen engine on Kubernetes with routing, metrics and autoscaling, and defend the choice using latency, throughput and cost per million tokens.

Six weeks. Twelve two-hour live sessions. Recordings, code, notes and practical materials included.

What you’ll learn

Benchmark, tune and deploy both vLLM and SGLang, then choose the right engine using measured latency, throughput and GPU cost.

  • Trace vLLM V1/PagedAttention and SGLang SRT/RadixAttention from HTTP request to generated tokens.

  • Build APIs with streaming, LoRA, structured output and tools; write SGLang programs with gen, select and fork.

  • Run one reproducible harness against both engines: TTFT, TPOT, inter-token latency and throughput under concurrent load.

  • Use torch profiler and Nsight to investigate bottlenecks before tuning.

  • Tune vLLM prefix caching and SGLang radix caching, overlap scheduling, chunked prefill, attention backends and CUDA graphs.

  • Produce FP8 and INT4 checkpoints with llm-compressor; compare quality, memory and latency on both engines against BF16.

  • Compare tensor, pipeline, data and expert parallelism on both engines, measuring multi-node communication and MoE scaling.

  • Measure EAGLE/MTP speculation and prefill/decode disaggregation using vLLM KV connectors and SGLang with Mooncake.

  • Deploy your chosen engine on Kubernetes with cache-aware routing, Prometheus metrics, autoscaling and failure handling.

  • Compare both engines on one workload; present p50/p99 latency, throughput, cost per million tokens and a migration plan.

Learn directly from expert instructors

Dr. Sreedath Panat

Dr. Sreedath Panat

MIT PhD and Vizuara co-founder teaching practical LLM systems.

Education & research
MIT
Dr. Rajat Dandekar

Dr. Rajat Dandekar

Vizuara Co-founder | Purdue PhD | Teaching engineers to build AI systems

Education & research
Purdue University
Dr. Raj Dandekar

Dr. Raj Dandekar

Vizuara co-founder | MIT PhD | GPU, inference and kernel engineering educator

Education, research & tools
MIT
The Julia Language
See all products from Rajat

Who this course is for

  • ML engineers who have run models on a GPU and want to benchmark, optimize and deploy reliable LLM serving systems.

  • Backend and infrastructure engineers building OpenAI-compatible endpoints, GPU deployments and production observability.

  • AI builders with Python and GPU experience who want to read, tune and compare the vLLM and SGLang codebases.

Prerequisites

  • Python, command-line fluency and GPU model experience

    You will read vLLM and SGLang code, run benchmarks and configure GPU servers. Inference fundamentals are covered in self-paced pre-work.

What's included

Live sessions

Learn directly from your instructors in a real-time, interactive format.

Lifetime access

Go back to course content and recordings whenever you need to.

Community of peers

Stay accountable and share insights with like-minded professionals.

Certificate of completion

Share your new skills with your employer or on LinkedIn.

Fundamentals pre-work

Self-paced inference fundamentals videos before week one, so live sessions focus on vLLM and SGLang code and experiments.

Reusable engineering toolkit

A shared benchmark harness, configs and deployment templates for both engines, plus lecture notebooks and cloud-GPU scripts.

Notes, assignments and quizzes

Handwritten notes, slides, assignments and interactive quizzes to reinforce the concepts behind each experiment.

Production capstone and showcase

Benchmark both engines, deploy your chosen endpoint, and present the SLO and cost analysis. Includes a capstone showcase page.

Maven Guarantee

Your purchase is backed by the Maven Guarantee.

Course syllabus

12 live sessions • 12 lessons

Week 1

Nov 10—Nov 15

    Inside both engines

    2 items

    Nov

    10

    Session 1

    Tue 11/103:30 AM—5:30 AM (UTC)

    Nov

    12

    Session 2

    Thu 11/123:30 AM—5:30 AM (UTC)

Week 2

Nov 16—Nov 22

    Serving on each engine

    2 items

    Nov

    17

    Session 3

    Tue 11/173:30 AM—5:30 AM (UTC)

    Nov

    19

    Session 4

    Thu 11/193:30 AM—5:30 AM (UTC)

Free resources

Schedule

Live sessions

4 hrs / week

Tuesdays and Thursdays, 9–11 AM IST, November 10–December 17, 2026. Twelve live lectures over six weeks, totaling 24 hours. Dr. Sreedath Panat leads the course, joined by Dr. Raj Dandekar and Dr. Rajat Dandekar. Recordings included.

    • Tue, Nov 10

      3:30 AM—5:30 AM (UTC)

    • Thu, Nov 12

      3:30 AM—5:30 AM (UTC)

    • Tue, Nov 17

      3:30 AM—5:30 AM (UTC)

Frequently asked questions

Maven for Teams

Reimbursement

Get your company to pay

Everything L&D needs: email template, receipts, and certificate of completion.

Get reimbursed

Team discount

Learn with your teammates

Save 20%+ when 2 or more teammates enroll in the same cohort.

Save 20%+ with a team

Private cohort

Run a cohort for your org

A dedicated cohort with a custom schedule and curriculum, tailored to your team.

Book a private cohort

$1,250

USD

Nov 10—Dec 17
Enroll