DeepSeek-V3: A Frontier Model Trained for $5.6M — How? | 5-Min Bite
Автор: Papers by Hand
Загружено: 2026-07-26
Просмотров: 75
Описание:
DeepSeek-V3 matched frontier models for a training bill of about $5.6M — a fraction of the usual cost. It's not one trick, it's four, stacked: a Mixture-of-Experts that wakes only 37B of its 671B parameters per token, a load-balancer with no auxiliary loss, FP8 training, and multi-token prediction. Here's how each one buys efficiency without giving up quality.
📄 Paper: DeepSeek-V3 Technical Report (arXiv:2412.19437)
⏱️ Chapters
0:00 A frontier model, trained cheap
0:48 DeepSeekMoE routing
1:27 Auxiliary-loss-free load balancing
2:37 FP8 training
3:38 Multi-token prediction & DualPipe
4:41 Results
5:23 Recap
This channel breaks down modern AI papers with original hand-drawn animations - the real mechanism, explained clearly, no fluff and no "OpenAI killed Claude and Claude killed Gemini BS"
👍 Like if this helped · 🔔 Subscribe for more paper breakdowns · 💬 Share your thoughts and requests in the comments.
#DeepSeek #AI #MachineLearning #LLM #DeepLearning
Повторяем попытку...
Доступные форматы для скачивания:
Скачать видео
-
Информация по загрузке: