ycliper

Популярное

Музыка Кино и Анимация Автомобили Животные Спорт Путешествия Игры Юмор

Интересные видео

2025 Сериалы Трейлеры Новости Как сделать Видеоуроки Diy своими руками

Топ запросов

смотреть а4 schoolboy runaway турецкий сериал смотреть мультфильмы эдисон
Скачать

Six Green Test Runs, Two Different Games: Testing Meta's Muse Spark 1.2

Автор: Agent Workflow Lab

Загружено: 2026-08-05

Просмотров: 96

Описание: Meta released Muse Code and Muse Spark 1.2 on 2026-08-05, and the developer blog claims the model "still generalizes to other coding agents you already use."

We tested that claim without installing Muse Code at all.

The task is Meta's own cookbook sample, bastion_breaker, at commit 9faa831c64de9ff461441a315b76aa1067cef940. It ships with one deliberately planted bug and a three-test oracle, and the unmodified sample fails exactly one test.

Because the shipped code labels its own bug in a comment, every model was run twice: once on Meta's file as shipped, and once with only those three comment lines deleted. The executable code is identical in both.

Three models were driven through a plain OpenRouter chat/completions call at temperature 0 — no agent loop, no tool calls, no Meta harness. Meta's oracle graded every result, and the tests were byte-compared afterward to confirm none had been weakened.

All six runs passed. Muse Spark 1.2 was the fastest and cheapest in the harder no-hint condition, at 5.88s and about $0.0119.

Then we asked something the oracle never asks: what happens when an enemy shot actually reaches the wall? The tests only assert that no bricks are destroyed. Two different games satisfy that. Four of the six passing patches let the shot through; two stopped it dead — and Muse Spark 1.2 produced both behaviours depending only on whether a non-executable comment was present.

Six green runs. Two different games.

Chapters
00:00 The claim
00:29 Meta's own broken sample
01:09 A comment that gives away the answer
01:41 No harness, just an API call
02:26 The question the oracle never asks
03:01 The same model, both behaviours
03:37 Which one is correct?
04:08 What this does and does not show
04:36 Green is not agreement

Scope and limits
One task, one prompt, one attempt per cell, six runs total. This is not a general coding benchmark and not a review of Muse Code, which was never installed or run. Costs are OpenRouter-reported for these specific calls. The probe covers one unspecified behaviour, not an exhaustive audit. A single run is never a model verdict.

Models tested
meta/muse-spark-1.2
anthropic/claude-opus-5
openai/gpt-5.6-sol

Source
Meta developer blog: https://developer.meta.com/ai/resourc...
Cookbook sample: https://github.com/meta-models/meta-m...

Agent Workflow Lab runs public, reproducible tests on AI agents and tools.
https://agent-workflow-lab.com/

Не удается загрузить Youtube-плеер. Проверьте блокировку Youtube в вашей сети.
Повторяем попытку...
Six Green Test Runs, Two Different Games: Testing Meta's Muse Spark 1.2

Поделиться в:

Доступные форматы для скачивания:

Скачать видео

  • Информация по загрузке:

Скачать аудио

Похожие видео

© 2025 ycliper. Все права защищены.



  • Контакты
  • О нас
  • Политика конфиденциальности



Контакты для правообладателей: [email protected]