vLLM 0.25: Model Runner V2 retires PagedAttention, full CUDA graphs explained
Автор: Learn AI Visually
Загружено: 2026-07-12
Просмотров: 49
Описание:
vLLM 0.25 removes PagedAttention, the KV-cache paging kernel that made vLLM famous, and makes Model Runner V2 the default for dense models.
PagedAttention stored the KV cache in scattered blocks and gathered them with a bespoke kernel. In 0.25 that standalone layer is gone, Model Runner V2 becomes the dense default, and vLLM reports full CUDA graphs: recording the whole decode step once and replaying it in a single launch instead of dispatching many tiny kernels live. This video explains what changed and why it matters, with the metaphor of a telephone switchboard replaced by an automatic exchange.
Full explainer (interactive): https://learnaivisually.com/g/vllm-0-...
Source: https://github.com/vllm-project/vllm/...
Learn AI and GPUs visually — free interactive courses at learnaivisually.com
#PagedAttention #vLLM #LLM #AI
Повторяем попытку...
Доступные форматы для скачивания:
Скачать видео
-
Информация по загрузке: