Agent · weight 5%智能体 · 权重 5%
WebArena
What it tests: A self-hosted replica of real websites (shopping, forums, maps, CMS). Agents must complete multi-step tasks by actually clicking and typing in a browser.
考什么:自托管的仿真网站环境(购物/论坛/地图/CMS),agent 要在浏览器里真点真输入,完成多步任务。
Why it counts: The standard for 'can this model operate a computer' — note: like all agent benchmarks, verify the setup; some reported scores use heavy scaffolding.
为什么算数:'模型会不会操作电脑'的标准考场。注意:Agent 类成绩普遍受脚手架影响大,我们只收录标注了测试方式的成绩。
| Maintainer维护方 | CMU |
|---|---|
| License许可 | Apache 2.0 |
| Reference point (prior-gen SOTA)参考点(上一代 SOTA) | 55 |
◈
Standings实时排名
11 models with sourced scores 个模型有溯源成绩| # | Model模型 | Score成绩 | Index指数 | Dated日期 | Measured by测评方 | |
|---|---|---|---|---|---|---|
| 1 | GPT-5.6 SolOpenAI | 92.2 | 100 | 2026-07-30 | 3rd-party第三方 | source ↗ |
| 2 | DeepSeek V3.2DeepSeek | 74.3 | 81 | 2026-02 | 3rd-party第三方 | source ↗×2 |
| 3 | Claude Opus 4.8Anthropic | 71.2 | 77 | — | 3rd-party第三方 | source ↗ |
| 4 | Gemini 3.1 Pro PreviewGoogle | 69 | 75 | — | 3rd-party第三方 | source ↗ |
| 5 | Muse Spark 1.1meta | 69 | 75 | — | 3rd-party第三方 | source ↗ |
| 6 | MiniMax M3MiniMax | 68.8 | 75 | — | 3rd-party第三方 | source ↗ |
| 7 | Claude Opus 5Anthropic | 68 | 74 | 2026-07-24 | vendor-reported厂商自报 | source ↗ |
| 8 | Claude Sonnet 5Anthropic | 64.5 | 70 | — | 3rd-party第三方 | source ↗×2 |
| 9 | MiniMax M2.7MiniMax | 63.1 | 68 | — | 3rd-party第三方 | source ↗ |
| 10 | GPT-5.5OpenAI | 59 | 64 | — | 3rd-party第三方 | source ↗×2⚠±3.7 |
| 11 | Grok 4.5xAI | 35 | 38 | — | 3rd-party第三方 | source ↗×2 |