Inside Cognition's inference stack: RL, speculative decoding & DFlash
Автор: Modal
Загружено: 2026-07-29
Просмотров: 729
Описание:
Modal x Cognition: Inside Devin's inference stack: RL, speculative decoding & DFlash
Modal's inference lead sits down with Cognition's research lead (Devin, Devin Desktop, Suite 1.6) for a deep dive into how RL and inference intersect, from training frontier coding models to serving them at scale.
They cover why 90% of an RL run is just execution, how Cognition uses agents to babysit training jobs overnight, and why inference is still an unsolved problem. The core of the conversation: DFlash, Modal's new diffusion-based speculator, and how it changes the speed/cost tradeoffs of production inference.
TIMESTAMPS:
0:00 — Intros
1:04 — What's hardest to get right in an RL run
3:39 — How Cognition started using Modal
4:03 — What to weigh when buying inference
4:58 — The latency/throughput Pareto frontier
7:14 — DFlash reveal (!!)
8:44 — Why faster runtime matters
9:56 — Memory-bound vs. compute-bound
10:57 — Tree-based speculation
14:20 — Hot take: RL and inference optimization converging
15:28 — Online training of the speculator
16:22 — "Auto inference"
17:27 — Agents competing on optimization problems
18:27 — What inference providers still get wrong
Повторяем попытку...
Доступные форматы для скачивания:
Скачать видео
-
Информация по загрузке: