Scratch to Scale: Real-World HPC Training

Zachary Mueller

Head of Developer Relations @ Lambda

This course is popular

7 people enrolled last week.

Learn Large-Scale AI Training on a Real Compute Cluster

Large-scale training has changed fast. Fine-tuning a model on one GPU does not prepare you for what happens when a training job has to run across many GPUs and machines.

At that scale, you have to understand how the whole system behaves. Communication between GPUs can become expensive. A poor choice of parallelism can waste memory or leave hardware idle. A job can look like it is running normally while a network bottleneck, an overloaded expert, or a slow worker cuts performance.

It is hard to learn this from documentation alone. Access to multi-node GPU clusters is limited, and many examples are small enough that the problems you see at scale never appear.

This course gives you access to a real cluster where you can work through those problems yourself. You will write distributed training code, inspect what happens across the cluster, measure where time and memory are going, and run workloads large enough for these decisions to matter. The course is built around the kind of work engineers face when training and running large models across many GPUs.

What you’ll learn

Build the distributed training skills to design, debug, and scale large AI workloads across real GPU clusters

  • Implement DDP and ZeRO-2 across multiple GPUs

  • Implement expert and context parallelism yourself

  • Combine DP, EP, and CP into real training jobs

  • Use Nsight and NCCL tools to inspect communications between GPUs

  • Find stalls, slow ranks, and other causes of poor utilization

  • Connect profiler traces to specific parts of your training code

  • Calculate memory use for parameters, gradients, optimizer state, and activations

  • Choose parallelism layouts that fit the model and hardware

  • Calculate MFU and HFU to measure how well the GPUs are being used

  • Reduce idle time by overlapping communication with computation

  • Use profiling data to decide which optimization is worth making

  • Launch and manage jobs with Slurm and Kubernetes-based infrastructure

  • Save and resume distributed jobs with sharded checkpoints

  • Train a 30B MoE workload on an HGX H100 node

  • Scale training to multiple HGX H100 nodes and larger model workloads

Learn directly from Zachary

Zachary Mueller

Zachary Mueller

I've worked in HPC for nearly a decade, starting from Hugging Face to now Lambda

Lambda
Hugging Face
See all products from Zach

Who this course is for

  • ML engineers comfortable with PyTorch who want hands-on experience training and debugging models across real GPU clusters.

  • AI infrastructure engineers who want to understand distributed training, parallelism, and performance beyond cluster setup.

  • Researchers and advanced practitioners who know model training but have had limited access to multi-node GPU systems.

What's included

Zachary Mueller

Live sessions

Learn directly from Zachary Mueller in a real-time, interactive format.

Real access to HPC Compute

Each student will get a chunk of a large HPC cluster throughout the entire cohort. No compute credits. No fighting for GPU availability.

Lifetime access

Go back to course content and recordings whenever you need to.

Generous office hours

Bring your blockers to office hours and leave with answers. Get feedback, debug help, and real support when you need it.

Community of peers

Stay accountable and share insights with like-minded professionals.

Certificate of completion

Share your new skills with your employer or on LinkedIn.

Course notebooks & code

Detailed course notebooks and material with meticulous notes to help walk you through the material and learn along the way

Maven Guarantee

Your purchase is backed by the Maven Guarantee.

Course syllabus

Week 1

Sep 14—Sep 20

    Sep

    15

    Course Introduction & Orchestration

    Tue 9/159:30 PM—11:00 PM (UTC)

    Sep

    17

    Accessing Compute, Interconnect Differences

    Thu 9/179:00 PM—10:00 PM (UTC)

Week 2

Sep 21—Sep 27

    Sep

    22

    DDP and ZeRO-2 from Scratch

    Tue 9/229:30 PM—11:30 PM (UTC)

    Sep

    24

    Debugging distributed communications

    Thu 9/249:00 PM—11:00 PM (UTC)

Schedule

Live sessions

1-3 hrs / week

    • Tue, Sep 15

      9:30 PM—11:00 PM (UTC)

    • Thu, Sep 17

      9:00 PM—10:00 PM (UTC)

    • Tue, Sep 22

      9:30 PM—11:30 PM (UTC)

Testimonials

  • Zach is my go to person on anything dealing with distributed training. He has maintained the most popular library in the world that helps developers with this problem, which means he’s familiar with all of the issues mere mortals have while tackling this problem. Zach is the best person to teach this subject. I am taking this course.
    Testimonial author image

    Hamel Husain

    Founder, Parlance Labs | Evals, evals, evals
  • Zach is one of the key people in the world making distributed machine learning more accessible. He has firsthand experience building some incredible popular tools like huggingface/accelerate. If you're GPU poor but considering moving to the GPU middle class then I can't think of a better instructor.
    Testimonial author image

    Mark Saroufim

    Software Engineer at Meta | Co-founder, GPU MODE
  • As a long time maintainer of HF Accelerate, Zach has had to master not only a deep understanding of ML scaling methods, but also to integrate them into a cohesive API for the masses to use. I've seen Zach consistently deliver robust, well-integrated solutions with a deep system-level understanding. You will be in good hands with Zach at the helm.
    Testimonial author image

    Stas Bekman

    Senior Machine Learning Engineer, Snowflake
  • Zach's stewardship of Accelerate and managing the intricacies of multiple distributed technologies (while abstracting it into an easy to use API) make Zach the preeminent leader in distributed training. Zach has shown deep understanding of everything from fundamentals to implementation, and is the first person that would come to mind to teach this
    Testimonial author image

    Wing Lian

    Founder, Axolotl
  • Zach is truly one in a million. I've never met anyone who puts so much time and thought into crafting deep learning code. With his background and experience, learning from him is an invaluable opportunity.
    Testimonial author image

    Radek Osmulski

    Senior Data Scientist, NVIDIA
  • Zach has a strong grasp of the fundamentals of fastai, but what really sets him apart is his ability to teach. He mixes in practical topics throughout his lessons, making every video engaging and worthwhile. With a proven track record of creating high-quality content, I’m confident that any course Zach produces will be worth your time and attention
    Testimonial author image

    Kevin Bird

    Co-Founder, Problem Solvers Guild

Frequently asked questions

Maven for Teams

Reimbursement

Get your company to pay

Everything L&D needs: email template, receipts, and certificate of completion.

Get reimbursed

Team discount

Learn with your teammates

Save 20%+ when 2 or more teammates enroll in the same cohort.

Save 20%+ with a team

Private cohort

Run a cohort for your org

A dedicated cohort with a custom schedule and curriculum, tailored to your team.

Book a private cohort

$5,000

USD

Sep 14Oct 31
Enroll