Flash Attention Explained — The Algorithm That Unlocked 128K Context Windows
Автор: AI deepdive
Загружено: 2026-05-31
Просмотров: 36
Описание:
Before 2022, a 128-thousand token context window was physically impossible. Then Flash Attention showed up. Two years later, models went from 2-thousand token contexts to over a million — on the same hardware.
In this video, we break down exactly how Flash Attention works, why it matters for local LLMs, and how it evolved from FA1 to FA3.
Timestamps:
0:00 — Hook: The Problem Flash Attention Solved
0:42 — The Quadratic Wall — Why Standard Attention Breaks
1:32 — GPU Memory Hierarchy — HBM vs SRAM
2:17 — IO-Awareness — The Key Insight
3:07 — Tiling — Processing Attention in Blocks
3:57 — Online Softmax — The Math Trick
4:52 — Flash Attention 1 vs 2 vs 3
5:52 — What Flash Attention Unlocked (Gemini, Claude, Llama)
6:42 — What It Means for Local LLMs
7:27 — Bottom Line
Subscribe to AI Deep Dive for more AI infrastructure explainers: / @aideepdive-x8i
#FlashAttention #AI #MachineLearning #LLM #DeepLearning #GPU #Transformer #AttentionMechanism #LocalLLM #PyTorch
Повторяем попытку...
Доступные форматы для скачивания:
Скачать видео
-
Информация по загрузке: