Kernel Engineering: From CUDA Foundations to Frontier AI Kernels

Dr. Raj Dandekar

Lead Instructor | MIT PhD

Dr. Sreedath Panat

MIT PhD | Vizuara Co-founder

+ Dr. Rajat Dandekar and Shubham Panchal

Build GPU kernels from first principles—and learn to explain their performance.

Move from using GPU libraries to understanding and writing the kernels behind them. Vizuara’s Kernel Engineering Workshop combines 25 live sessions with hands-on coding, an illustrated book and a graded capstone.

Lead instructor Dr. Raj Dandekar teaches concepts and the capstone; Shubham Panchal leads live coding and frontier deep-dives. Start with CPU parallelism, roofline reasoning, CUDA and memory hierarchy. Build GEMM kernels, use tensor cores, profile bottlenecks and implement FlashAttention.

Then explore Hopper, Triton, CUTLASS/CuTe, inference-serving kernels, Blackwell/NVFP4, Flash Attention 4, DeepSeek’s FlashMLA/DeepGEMM and the agent-plus-profiler loop for AI-written kernels.

The workshop runs in partnership with Crusoe, with a guest lecture and capstone project ideas. Present your completed kernel project at a graded demo day.

For learners comfortable with Python and basic PyTorch; no prior CUDA or C++ is assumed. Cloud-GPU setup is covered. The website lists October 12–December 7, 2026, with recordings and code included.

What you’ll learn

Write, profile and optimize GPU kernels, then apply modern kernel techniques in a graded engineering capstone.

  • Use roofline reasoning, CUDA execution and memory hierarchy to explain performance limits.

  • Build transpose and GEMM kernels; apply tiling, vectorization, warp tiling and tensor cores.

  • Use Nsight Compute and compiled instructions to profile and debug kernels.

  • Build FlashAttention and compare the architectural ideas behind FA2, FA3 and FA4.

  • Explore Hopper, Blackwell, Triton, CUTLASS/CuTe, low precision and inference-serving kernels.

  • Study DeepSeek kernels and evaluate AI-generated kernels with an agent-plus-profiler loop.

  • Work on a kernel problem drawn from project ideas shared by Crusoe.

  • Apply techniques from the six-part curriculum and present your work at a graded demo day.

  • Use the included illustrated book, code and interview-preparation website to consolidate your learning.

Learn directly from expert instructors

Dr. Raj Dandekar

Dr. Raj Dandekar

Vizuara Co-founder | MIT PhD | Kernel concepts, advanced lectures & capstone

Education, research & tools
MIT
The Julia Language
Dr. Sreedath Panat

Dr. Sreedath Panat

Vizuara Co-founder | MIT PhD | Computer vision & scientific machine learning

Education & research
MIT
Dr. Rajat Dandekar

Dr. Rajat Dandekar

Vizuara Co-founder | Purdue PhD | Teaching engineers to build AI systems

Education & research
Purdue University
Shubham Panchal

Shubham Panchal

On-device ML engineer | SmolChat creator | Live coding & frontier deep-dives

See all products from Rajat

Who this course is for

  • Engineers comfortable with Python and basic PyTorch who want to write GPU kernels. No previous CUDA or C++ experience is required.

  • ML and inference engineers seeking hands-on practice with GEMM, attention, profiling and modern GPU kernel toolchains.

What's included

Live sessions

Learn directly from your instructors in a real-time, interactive format.

All code files and session recordings

Keep the code from the workshop and revisit recordings of all live sessions. Live participation is encouraged, especially for the coding sessions.

Kernel Engineering Interview Preparation

Access Vizuara’s Kernel Engineering Interview Preparation website alongside the workshop.

Illustrated Kernel Engineering Book

Full access to Vizuara’s illustrated Kernel Engineering Book, the written companion to the live sessions, unlocked at registration.

Crusoe collaboration and graded capstone

A Crusoe guest lecture and capstone project ideas, followed by a graded demo day. The demo-day date is announced inside the cohort.

One month of Vizz-AI and Vizuara AI Pods

The workshop includes one month of access to Vizz-AI and Vizuara AI Pods as a bonus, as listed on the workshop website.

Maven Guarantee

Your purchase is backed by the Maven Guarantee.

Course syllabus

Week 1

Oct 12—Oct 18

    Week 1 — CPU parallelism and the roofline model

    3 items

Week 2

Oct 19—Oct 25

    Week 2 — CUDA programming and memory

    3 items

Schedule

Live sessions

50 hrs

25 live sessions, two hours each, October 12–December 7, 2026. Classes run Monday, Wednesday and Friday, 7–9 AM India Standard Time (UTC+5:30). Recordings are included; calendar invitations and meeting links will be shared with enrolled students.

Vizuara: first principles, working code, proven teaching

Vizuara brings first-principles explanations, live coding, research papers and hands-on projects to a global AI learning community. Our YouTube channel has 224,000 subscribers, and our platform has more than 19,000 registered accounts across school, institutional and professional programs.

Founded by Purdue PhD Dr. Rajat Dandekar and MIT PhDs Dr. Raj Dandekar and Dr. Sreedath Panat, all IIT Madras alumni. The founders co-authored Manning’s Build a DeepSeek Model (From Scratch).

Vizuara’s collection of 109 learner stories reports a 4.96/5 average across rated reviews, with reviewers from organizations including Bosch, ISRO and Confluent. This feedback comes from earlier Vizuara programs.

Read learner stories: https://reviews.vizuara.ai/

Watch our teaching: https://www.youtube.com/@vizuara

Free lectures — coming soon

5D Parallelism — free lecture coming soon.

Inference Engineering — free lecture coming soon.

We are preparing the lecture links for this resource collection. Recordings are not available here yet; this section will be updated when they are ready.

Frequently asked questions

Maven for Teams

Reimbursement

Get your company to pay

Everything L&D needs: email template, receipts, and certificate of completion.

Get reimbursed

Team discount

Learn with your teammates

Save 20%+ when 2 or more teammates enroll in the same cohort.

Save 20%+ with a team

Private cohort

Run a cohort for your org

A dedicated cohort with a custom schedule and curriculum, tailored to your team.

Book a private cohort

$3,000

USD

Oct 12Dec 7
Enroll