Fast Inference of Mixture-of-Experts Language Models with Offloading
Автор: AI Papers Academy
Загружено: 2024-01-12
Просмотров: 2146
Описание:
In this video we review a recent important paper titled: "Fast Inference of Mixture-of-Experts Language Models with Offloading".
Mixture of Experts (MoE) is an important strategy to improve the efficiency of transformer based large language models (LLMs) nowadays.
However, MoE models usually have a large memory footprint since we need to load the weights of all experts. This makes it hard to run MoE models on low tier GPUs.
This paper introduces a method to efficiently run transformer based MoE LLMs on a limited memory environment using offloading techniques. Specifically, the researchers are able to run Mixtral-8x7B on the free-tier version of Google Colab.
In the video, we provide a reminder for how mixture of experts works, and then dive into the offloading method presented in this paper.
-----------------------------------------------------------------------------------------------
Paper page - https://arxiv.org/abs/2312.17238
Soft MoE - • Soft Mixture of Experts - An Efficient Spa...
Code - https://github.com/dvmazur/mixtral-of...
Post - https://aipapersacademy.com/moe-offlo...
Original Mixture-of-Experts paper review - https://aipapersacademy.com/mixture-o...
-----------------------------------------------------------------------------------------------
✉️ Join the newsletter - https://aipapersacademy.com/newsletter/
👍 Please like & subscribe if you enjoy this content
We use VideoScribe to edit our videos - https://tidd.ly/44TZEiX (affiliate)
-----------------------------------------------------------------------------------------------
Chapters:
0:00 Paper Introduction
1:34 Mixture of Experts
3:44 MoE Offloading
10:29 Mixed MoE Quantization
11:13 Inference Speed
Повторяем попытку...
Доступные форматы для скачивания:
Скачать видео
-
Информация по загрузке: