ycliper

Популярное

Музыка Кино и Анимация Автомобили Животные Спорт Путешествия Игры Юмор

Интересные видео

2025 Сериалы Трейлеры Новости Как сделать Видеоуроки Diy своими руками

Топ запросов

смотреть а4 schoolboy runaway турецкий сериал смотреть мультфильмы эдисон
Скачать

Streaming Massive AI Models Directly from an NVMe SSD

Автор: BoredBrains Consortium

Загружено: 2026-08-12

Просмотров: 59

Описание: Who says you need a massive GPU cluster to run State-of-the-Art open-weight models?
In this video, I’m showing off a brand new AI execution engine I built from scratch in Rust. It’s less than 200 Kilobytes in size, but it allows me to run a massive 72-Billion parameter model (Qwen 72B), alongside a Gemma 7B and a TinyLlama 1.1B, all at the exact same time on my local rig.
How does it work?
I bypassed standard Python/PyTorch wrappers. This Rust library dynamically cascades the AI's weight matrices. It fills up my dual RTX 3090s, overflows into the 128GB of system RAM, and when that fills up, it streams the rest of the model directly off the NVMe SSD!
What you’ll see in this demo:
🚀 The EasyBake AI App: A quick update that the app is officially LIVE on the Google Play Store (Grab it for $5 before the subscription model kicks in!).
🧠 The Rust Engine: Why I threw out legacy AI frameworks to build something that runs natively on PolyMorphOS.
💾 NVMe Streaming: Watching the 1.1B model run flawlessly while streaming purely from the hard drive (bypassing the GPU entirely).
💥 Model Fragility: A candid look at how fragile standard SOTA models are compared to our custom JARVITS architectures when you push them too hard.
📊 The B-Top Telemetry: Watching the system effortlessly balance 94GB of VRAM and 91GB of system RAM across the CPU and GPU.
The hyperscalers want you to stay dependent on their cloud. We are building the tools to put that power back on your desk.
Drop a comment below: What massive model would you run if your hard drive could act as RAM?
[Chapters]
0:00 - EasyBake AI is LIVE on the Play Store!
1:30 - Replacing PyTorch with a 200KB Rust Engine
3:15 - Loading Qwen 72B, Gemma 7B, and TinyLlama
4:45 - Model Fragility (Why JARVITS is tougher)
6:10 - Software-Defined Unified RAM (VRAM to SSD)
8:00 - Booting TinyLlama directly off the NVMe Drive
9:50 - The 72B Model Hydration and Telemetry
11:30 - CPU & GPU Working in Unison
13:15 - You Don't Need a Data Center!
🌐 Get the EasyBake AI App here: [https://play.google.com/store/apps/de...

#EdgeAI #Rust #MachineLearning #AI #LocalAI #TechDemo #SoftwareEngineering #PolyMorphOS

Не удается загрузить Youtube-плеер. Проверьте блокировку Youtube в вашей сети.
Повторяем попытку...
Streaming Massive AI Models Directly from an NVMe SSD

Поделиться в:

Доступные форматы для скачивания:

Скачать видео

  • Информация по загрузке:

Скачать аудио

Похожие видео

© 2025 ycliper. Все права защищены.



  • Контакты
  • О нас
  • Политика конфиденциальности



Контакты для правообладателей: [email protected]