Qwen3.8-Max-0902’s Terminal Score Jumped 157% — Same Price
Автор: Morgans Code
Загружено: 2026-09-02
Просмотров: 3716
Описание:
Qwen3.8-Max-0902 just pushed its Terminal-Bench 3.0 result from 11.3% to 29% — roughly a 2.6× jump — without increasing the model size, 1M-token context window or standard API price.
On DeepSWE 1.1, Alibaba reports a score of 69.3, reaching roughly 95% of GPT-5.6 Sol’s published result. The updated model also debuted at #1 in Code Arena’s WebDev leaderboard with 1,691 points.
So what changed inside Qwen3.8-Max-0902?
Instead of training a larger foundation model, Qwen applied another round of post-training focused on coding and “Cowork” workflows. That means longer agent tasks where the model must inspect files, use tools, run commands, verify results, recover from errors and continue working until the job is finished.
In this video, we examine whether Qwen’s 0902 update represents a genuine leap for AI coding agents — and what it could reveal about the upcoming Qwen 4 generation.
We cover:
• Qwen3.8-Max-0902 vs the original Qwen3.8-Max
• Terminal-Bench 3.0: 11.3% to 29%
• Why a 2.6× terminal coding jump matters
• DeepSWE 1.1: Qwen 0902 vs GPT-5.6 Sol
• How much of the old Qwen–Sol gap was closed
• Qwen’s improvements across eight coding evaluations
• Code Arena’s WebDev leaderboard result
• Where Qwen now beats Claude Opus 5
• Why coding-focused post-training can create such large gains
• Qwen3.8-Max-0902 API pricing and 1M-token context
• What this update could tell us about Qwen 4
According to Qwen’s published results, all eight coding evaluations improved. Terminal-Bench produced the largest relative gain, while QwenSWEbench V2, JobBench and DeepSWE also moved significantly higher.
The external signal is Code Arena. Qwen3.8-Max-0902 debuted at #1 with 1,691 points, ahead of Claude Opus 5 Max, Kimi K3 Max and the original Qwen3.8-Max. Code Arena scores and ranks are live and preliminary, so they may change as more votes and new models arrive.
Important context:
• Most coding benchmark numbers discussed here come from Qwen’s own release table.
• Code Arena is an external evaluation, but its scores and confidence intervals continue to update.
• Strong benchmark results do not guarantee better performance on every real repository or coding-agent setup.
• Qwen3.8-Max-0902 is currently a hosted proprietary API model, not a separately released open-weight checkpoint.
• Token pricing does not directly measure the total cost of completing a real software task.
The standard QwenCloud price remains $2 per million input tokens and $6 per million output tokens. GPT-5.6 Sol is more expensive per token, but a proper cost comparison also requires measuring retries, output length and successful task completion.
The next useful test is straightforward: run the original Qwen3.8-Max, Qwen3.8-Max-0902 and GPT-5.6 Sol on the same real repository with the same coding agent.
Which comparison would you like to see next: Qwen 0902 vs GPT-5.6 Sol on a real project, or a full breakdown of what 0902 could mean for Qwen 4?
Chapters:
00:00 Qwen3.8-Max-0902: What Changed?
00:54 Same Model, Stronger Post-Training
01:52 Terminal-Bench 3.0: From 11.3% to 29%
02:40 Every Qwen Coding Benchmark Improved
03:19 Qwen Debuts at #1 in Code Arena
04:10 Qwen 0902 vs GPT-5.6 Sol
05:13 Qwen vs Claude Opus 5 Coding Tests
05:53 Why Post-Training Changed Everything
06:41 Same Model, Same API Price
07:19 What Real-World Testing Still Needs
08:09 What Qwen 0902 Reveals About Qwen 4
Watch next:
• Qwen 3.8 27B Is Coming — Sonnet 5 on Your Laptop?
• Qwen 3.8 27B Is Coming — Sonnet 5 on Your ...
• Every Major AI Lab Has a Big Release Coming
• Every Major AI Lab Has a Big Release Coming
Sources:
• Qwen3.8-Max-0902 announcement:
https://x.com/Alibaba_Qwen/status/209...
• QwenCloud model specifications and pricing:
https://www.qwencloud.com/models/qwen...
• Original Qwen3.8 release:
https://qwen.ai/blog?id=qwen3.8
• Code Arena WebDev leaderboard:
https://arena.ai/leaderboard/code
• Terminal-Bench:
https://www.tbench.ai/benchmarks
• GPT-5.6 Sol DeepSWE result:
https://www.together.ai/blog/deepseek...
Narration voice: VCTK speaker p360 — CSTR VCTK Corpus 0.92, University of Edinburgh (CC BY 4.0).
https://datashare.ed.ac.uk/handle/102...
Voice synthesis: Chatterbox by Resemble AI (MIT License).
https://github.com/resemble-ai/chatte...
#Qwen #AICoding #CodingAgents
Повторяем попытку...
Доступные форматы для скачивания:
Скачать видео
-
Информация по загрузке: