Distributed Mixture of Experts MoE Visualised DeepSeek
Автор: HowCanAIHelp
Загружено: 2026-05-15
Просмотров: 204
Описание:
In this video we move beyond the basic math and visualise exactly how the Mixture of Expert forward pass is scaled across the GPU cluster. If you want to understand how frontier Large Language Models (LLMs) actually run mixture of experts on physical hardware, this is the video for you.
In this walkthrough I visualize DeepSeek V3 cluster for explaining:
🔹 Routed Experts & All-Reduce: How distributed GPUs calculate partial expert outputs and synchronize them using the Ring All-Reduce collective operation.
🔹 Tensor Parallelism (TP): How massive weight matrices are physically sliced using Column and Row Parallelization.
🔹 The Shared Expert: The step-by-step distributed calculation of parallelised expert
🔹 The Distributed Cluster: Bringing it all together to see how the Shared Expert and Routed Experts run simultaneously across a multi-node supercomputer.
You can watch Part 1 to understand how shared and routed expert and gate are implemented: • Mixture of Experts (MoE) Visually Explaine...
⏱️ Chapters:
00:00 - Cluster Overview
00:26 - How Routed Experts are Calculated
01:30 - The All-Reduce Operation Visualized
05:34 - Tensor Parallelism (Row & Column Slicing)
07:42 - Distributed Expert
08:43 - Running the Shared Expert on a Distributed Cluster
Whether you are an ML engineer, a computer science student, or just fascinated by distributed computing, this step-by-step visual guide will help you master the hardest concepts in AI scaling.
Tags: #DeepSeek #MixtureOfExperts #DistributedComputing #MachineLearning #TensorParallelism #PyTorch #ArtificialIntelligence #LLM #GPU #DeepLearning #SoftwareEngineering #AllReduce
In this video I visualise how Mixture of Experts (MoE) is implemented in the DeepSeek V3 model.
The walkthrough material is based on open source inference code and technical report:
https://github.com/deepseek-ai/DeepSe...
https://arxiv.org/pdf/2412.19437
Link to GLU Variants Improve Transformer by Noam Shazeer:
https://arxiv.org/pdf/2002.05202
In this visual walkthrough, I cover:
0:00 - MoE Architecture Components
2:58 - Expert Breakdown
4:21 - Studying the Gate
5:08 - Routing to specialised experts
Whether you are a PyTorch developer, an AI researcher, or just a tech enthusiast trying to understand how modern Large Language Models (LLMs) operate, this visual breakdown will help you build a true intuition for MoE architecture.
🔜 Want to know how this actually runs on a distributed cluster? Subscribe so you don't miss Part 2, where I cover how Deepseek calculates Mixture of Expert with Tensor Parallelism and All Reduce on a 16 GPUs.
Tags:
#DeepSeek #MixtureOfExperts #MachineLearning #AI #ArtificialIntelligence #LLM #PyTorch #DeepLearning #NeuralNetworks #SoftwareEngineering #TransformerModel
Повторяем попытку...
Доступные форматы для скачивания:
Скачать видео
-
Информация по загрузке: