StreamingLLM Lecture
Автор: MIT HAN Lab
Загружено: 2023-10-24
Просмотров: 3827
Описание: Streaming Language Models with Attention Sinks: deploying LLMs for streaming applications with long text sequences using limited memory poses significant challenges. We find existing window-based KV cache adopts a suboptimal KV cache eviction policy. We unveil the "attention sink" phenomenon where initial tokens receives strong attention, and should never be evicted from the KV cache. Leveraging this, we introduce StreamingLLM, which always keeps the attention sinks in the KV cache and the rest in a sliding window mechanism, enabling LLMs to process infinite text lengths without fine-tuning. Code: https://github.com/mit-han-lab/stream...
Повторяем попытку...
Доступные форматы для скачивания:
Скачать видео
-
Информация по загрузке: