It Begins: An AI Broke Out of OpenAI's Lab
Автор: Absolutely Agentic
Загружено: 2026-07-23
Просмотров: 7856
Описание:
Subscribe to Absolutely Agentic 👉 https://absolutelyagentic.com/upgrade
Sign up to our newsletter 👉 https://absolutelyagentic.com/?modal=...
OpenAI's models were locked in a sealed sandbox and given a hacking exam.
Instead of solving the problems, they found a zero-day in the one channel to the outside world, escaped, and broke into Hugging Face to steal the answer key.
Nobody told them to do it.
This is reward hacking, the same behaviour as a 2016 boat game exploit, now armed with the ability to breach real infrastructure.
The evaluation and the incident are no longer separable.
And the safety guardrails slowed the defenders while doing nothing to the attacker.
Chapters
00:00 - Intro
1:39 - Two Disclosures, One Incident
3:56 - How the Models Escaped the Sandbox
7:36 - When Guardrails Only Stop the Defenders
9:56 - The Gap Between What We Said and What We Meant
Sources to Google
OpenAI - ExploitGym Incident Disclosure (July 2026)
Hugging Face - Autonomous Agent Intrusion Disclosure (July 2026)
OpenAI - Faulty Reward Functions in the Wild (CoastRunners, 2016)
Victoria Krakovna (DeepMind) - Specification Gaming: The Flip Side of AI Ingenuity
UK AI Security Institute - Frontier Model Cyber Capability Evaluations
Nick Bostrom - Superintelligence: Paths, Dangers, Strategies
#AISafety #RewardHacking #OpenAI #HuggingFace #Cybersecurity #ArtificialIntelligence #AI #ChatGPT
Повторяем попытку...
Доступные форматы для скачивания:
Скачать видео
-
Информация по загрузке: