Mechanics of Self-Attention Next Token Prediction
Автор: Knut Jägersberg
Загружено: 2025-09-25
Просмотров: 28
Описание: The academic paper investigates the mechanics of next-token prediction in Transformer-based language models, focusing specifically on how a single self-attention layer learns this objective using gradient descent. The authors demonstrate that this process implicitly discovers a two-step automaton: hard retrieval to select high-priority input tokens, followed by soft composition to generate the next token as a convex combination of those selected. This behavior is formally characterized by linking the learned attention weights to a Support Vector Machine (SVM) formulation defined by token-priority graphs (TPGs) and their strongly-connected components (SCCs), which encode the sequential priority orders in the training data. The analysis, supported by theorems and experiments, shows that the attention weights decompose into an unbounded component (hard retrieval) and a finite component (soft composition).
Повторяем попытку...
Доступные форматы для скачивания:
Скачать видео
-
Информация по загрузке: