ycliper

Популярное

Музыка Кино и Анимация Автомобили Животные Спорт Путешествия Игры Юмор

Интересные видео

2025 Сериалы Трейлеры Новости Как сделать Видеоуроки Diy своими руками

Топ запросов

смотреть а4 schoolboy runaway турецкий сериал смотреть мультфильмы эдисон
Скачать

LLM Inference Optimization Explained — From 8 Tokens/sec to 50+

Автор: AI deepdive

Загружено: 2026-06-13

Просмотров: 80

Описание: Why does a 70B language model crawl at 8 tokens per second on one setup, then feel instant on another? The difference is inference optimization: KV cache management, PagedAttention, continuous batching, quantization, speculative decoding, model parallelism, and production serving frameworks like vLLM and TensorRT-LLM.

In this AI Deep Dive, we break down the systems engineering behind fast LLM serving — the techniques that turn expensive, slow autoregressive generation into real-time user experiences.

Timestamps:
0:00 — Hook: The Inference Optimization Gap
0:45 — KV Cache: The Bottleneck Behind LLM Serving
2:00 — PagedAttention: Virtual Memory for Attention
3:10 — Continuous Batching: Keeping GPUs Full
4:20 — Quantization: Shrinking Model Weights
5:30 — Speculative Decoding: Parallelizing Token Generation
6:25 — Model Parallelism: Splitting Giant Models Across GPUs
7:20 — Serving Frameworks: vLLM vs TensorRT-LLM
8:20 — Bottom Line: Good Engineering Applied to the Right Bottlenecks

Subscribe to AI Deep Dive for more AI infrastructure explainers:    / @aideepdive-x8i  

#LLM #InferenceOptimization #AIInfrastructure #vLLM #TensorRTLLM #PagedAttention #Quantization #SpeculativeDecoding #LocalLLM #MachineLearning

Не удается загрузить Youtube-плеер. Проверьте блокировку Youtube в вашей сети.
Повторяем попытку...
LLM Inference Optimization Explained — From 8 Tokens/sec to 50+

Поделиться в:

Доступные форматы для скачивания:

Скачать видео

  • Информация по загрузке:

Скачать аудио

Похожие видео

© 2025 ycliper. Все права защищены.



  • Контакты
  • О нас
  • Политика конфиденциальности



Контакты для правообладателей: [email protected]