The AI tools index that doesn't waste your time.不浪费你时间的 AI 工具索引。
Agent · weight 5%智能体 · 权重 5%

WebArena

official site官网

What it tests: A self-hosted replica of real websites (shopping, forums, maps, CMS). Agents must complete multi-step tasks by actually clicking and typing in a browser.

考什么:自托管的仿真网站环境(购物/论坛/地图/CMS),agent 要在浏览器里真点真输入,完成多步任务。

Why it counts: The standard for 'can this model operate a computer' — note: like all agent benchmarks, verify the setup; some reported scores use heavy scaffolding.

为什么算数:'模型会不会操作电脑'的标准考场。注意:Agent 类成绩普遍受脚手架影响大,我们只收录标注了测试方式的成绩。

Maintainer维护方CMU
License许可Apache 2.0
Reference point (prior-gen SOTA)参考点(上一代 SOTA)55

Standings实时排名

11 models with sourced scores 个模型有溯源成绩
#Model模型Score成绩Index指数Dated日期Measured by测评方
1GPT-5.6 SolOpenAI92.21002026-07-303rd-party第三方source ↗
2DeepSeek V3.2DeepSeek74.3812026-023rd-party第三方source ↗×2
3Claude Opus 4.8Anthropic71.2773rd-party第三方source ↗
4Gemini 3.1 Pro PreviewGoogle69753rd-party第三方source ↗
5Muse Spark 1.1meta69753rd-party第三方source ↗
6MiniMax M3MiniMax68.8753rd-party第三方source ↗
7Claude Opus 5Anthropic68742026-07-24vendor-reported厂商自报source ↗
8Claude Sonnet 5Anthropic64.5703rd-party第三方source ↗×2
9MiniMax M2.7MiniMax63.1683rd-party第三方source ↗
10GPT-5.5OpenAI59643rd-party第三方source ↗×2⚠±3.7
11Grok 4.5xAI35383rd-party第三方source ↗×2

How scores become the Intelligence Score →成绩如何合成智能评分 →