GLM-5.3 explained in 8 minutes
Автор: Gork Explained
Загружено: 2026-08-18
Просмотров: 19
Описание:
GLM-5.3 scores 28.3 on Terminal-Bench 3.0 where GLM-5.2 scored 4.6, and the base model underneath the two of them is exactly the same 743 billion parameters with the same 39 billion active per token. No new layers, no wider context window, no second pre-training run. The whole jump came out of post-training, and out of one idea Z.ai calls environment scaling: instead of feeding the model more text, you drop it into a very large number of simulated working environments and let it act, fail and get a signal back.
This video draws that mechanism: what stayed frozen, what pre-training and post-training each give a model, what an environment actually is, the loop that runs inside one, and why a model that learns no new fact still climbs 6x on a benchmark made of long multi-step terminal tasks. It ends on the security results Z.ai says it was not aiming for, and on what a longer chain is worth to somebody who hands a ticket to an agent and walks away.
Every figure in the video is Z.ai's own, published on the 14th of August 2026, and nobody outside the company has reproduced any of them yet. The video says so twice.
Figures drawn: Terminal-Bench 3.0 4.6 to 28.3 · DeepSWE v1.1 46.2 to 66.9 · CyberGym 77.2% to 84.5% · ExploitBench 24.4% to 54.4% · AutomationBench 48.2% · GDPval-AA v2 1769 Elo · internal code bench 31.4% at ~50,000 output tokens per task · 2,436 vulnerabilities across 269 open source projects.
Повторяем попытку...
Доступные форматы для скачивания:
Скачать видео
-
Информация по загрузке: