Introducing the Artificial Analysis Intelligence Index v4.1
Автор: Artificial Analysis
Загружено: 2026-06-15
Просмотров: 446
Описание:
The Artificial Analysis Intelligence Index is our synthesis metric for assessing model intelligence and tracking AI progress. Today we announced v4.1 with changes listed below.
Links:
Artificial Analysis Intelligence Index - https://artificialanalysis.ai/
Artificial Analysis Intelligence Index Methodology - https://artificialanalysis.ai/methodo...
Intelligence Index v4.1 brings the following updates:
1. We upgraded three evaluations, removed one, and reweighted the Intelligence Index:
➤ Upgraded Terminal-Bench Hard to Terminal-Bench 2.1 and τ²-Bench Telecom to τ³-Bench Banking. Both move to newer, more robust task sets with harder, more realistic agentic scenarios that better separate frontier models
➤ Upgraded GDPval-AA to GDPval-AA v2. The upgrade re-baselines Elo to human performance at 1000, introduces a rotating panel of frontier-model judges, and raises the turn limit from 100 to 250 for longer-horizon agent trajectories
➤ Removed IFBench due to saturation. The benchmark no longer distinguishes frontier models sufficiently, so we have removed it from the Intelligence Index. We will continue to run it and publish results on new model releases
2. Cost per Task, Time per Task, and Tokens per Task:
Three new per-task metrics, reported for every model and based on the Intelligence Index. We take the total cost, total time, and total output tokens for a model to run the Intelligence Index and divide by the number of tasks across its evaluations, giving the average cost, time, and output tokens to complete a single Intelligence Index task
3. Cached input token reporting:
We now report cached input tokens and their impact on cost, including the cost to run the Intelligence Index, to better reflect the real cost of running each model
Timestamps:
0:00 - Intro
0:26 - Evaluation and Benchmark Changes
1:00 - Cost per Task, Time per Task, Tokens per Task additional metrics
1:39 - Cache pricing breakdown
1:54 - Leading overall models
2:05 - Leading open-source models
2:19 - Methodology and outro
Повторяем попытку...
Доступные форматы для скачивания:
Скачать видео
-
Информация по загрузке: