NVIDIA fixed a FLAW in LINEAR ATTENTION nobody was talking about (Gated DeltaNet-2)
Автор: Machine Learning With Aaryan
Загружено: 2026-05-23
Просмотров: 121
Описание:
NVIDIA spotted a constraint hiding inside linear attention that nobody was talking about — and fixed it in Gated DeltaNet-2.
In this video I explain everything from scratch:
→ why standard attention explodes quadratically
→ how Gated DeltaNet replaces it with a fixed-size notepad
→ the delta rule — writing corrections instead of accumulating blindly
→ the hidden scalar constraint in GDN that chained erasing and writing together
→ how GDN-2 decouples them into two independent vectors
This is also the architecture behind Qwen3.6 and 3.7 models - used in a 3:1 ratio with full attention layers.
━━━━━━━━━━━━━━━━━━━━━━━━
CHAPTERS
━━━━━━━━━━━━━━━━━━━━━━━━
0:00 Intro
0:23 Standard attention — the quadratic problem
0:53 How Gated DeltaNet is different
1:08 The notepad — fixed-size memory
1:38 The delta rule — write only the correction
2:05 Forgetting — the gate that fades old memory
2:37 The hidden constraint in GDN
3:16 GDN-2 — two independent vectors
4:00 Architecture walkthrough
4:34 The lineage: DeltaNet → GDN → KDA → GDN-2
5:36 Wrap up
━━━━━━━━━━━━━━━━━━━━━━━━
PAPER
━━━━━━━━━━━━━━━━━━━━━━━━
Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention
Hatamizadeh, Choi, Kautz · NVIDIA · May 2026
arXiv: https://arxiv.org/abs/2605.22791
GitHub: https://github.com/NVlabs/GatedDeltaN...
━━━━━━━━━━━━━━━━━━━━━━━━
TAGS
━━━━━━━━━━━━━━━━━━━━━━━━
#GatedDeltaNet #GatedDeltaNet2 #LinearAttention #NVIDIA #Qwen3 #LLM #AIResearch #DeepLearning #MachineLearning #Transformers
Повторяем попытку...
Доступные форматы для скачивания:
Скачать видео
-
Информация по загрузке: