The AI tools index that doesn't waste your time.不浪费你时间的 AI 工具索引。
Preference · weight 5%偏好 · 权重 5%

LMArena (Chatbot Arena)

official site官网

What it tests: Millions of blind side-by-side human votes: users compare two anonymous models and pick the better answer. Scores are Bradley-Terry ratings (Elo-like).

考什么:数百万次真人盲测投票:用户并排看两个匿名模型的回答,选更好的那个。分数是 Bradley-Terry 评分(类 Elo)。

Why it counts: The only large-scale measure of what humans actually prefer — but 2025's 'Leaderboard Illusion' paper showed big labs can game it, so we cap its weight at 5%.

为什么算数:唯一大规模反映'真人真实偏好'的数据——但 2025 年《排行榜幻觉》论文实锤了大厂能刷榜,所以权重压到 5% 并保留争议说明。

Maintainer维护方Arena AI, Inc.(原 UC Berkeley LMSYS)
License许可CC-BY-4.0(官方数据集)
Reference point (prior-gen SOTA)参考点(上一代 SOTA)1450

Standings实时排名

48 models with sourced scores 个模型有溯源成绩
#Model模型Score成绩Index指数Dated日期Measured by测评方
1Claude Opus 5Anthropic1699⚡max1002026-08-05official官方榜source ↗×8⚠±194.25
2GPT-5.6 SolOpenAI1623⚡max812026-07-31official官方榜source ↗×10⚠±142.2
3GPT-5.6 TerraOpenAI1522⚡max562026-07-31official官方榜source ↗×5⚠±57
4Claude Fable 5Anthropic1506522026-08-12official官方榜source ↗×9
5Muse Spark 1.2meta1498502026-08-12official官方榜source ↗
6Qwen3.8 Max (0902)Qwen1491482026-08-12official官方榜source ↗×3⚠±367
7Kimi K3Moonshot AI1489⚡max482026-08-12official官方榜source ↗×7⚠±193
8Muse Spark 1.1meta1489482026-08-12official官方榜source ↗
9Gemini 3.1 Pro PreviewGoogle1486472026-08-12official官方榜source ↗×9
10Gemini 3.6 FlashGoogle1484⚡high462026-08-12official官方榜source ↗×6⚠±64
11GPT-5.5OpenAI1482⚡high462026-08-12official官方榜source ↗×4
12GPT-5.6 Sol ProOpenAI1481⚡max462026-08-12official官方榜source ↗
13Gemini 3.5 FlashGoogle1477452026-08-12official官方榜source ↗×7
14Qwen3.7 FlashQwen1475443rd-party第三方source ↗
15Grok 4.20xAI1475442026-08-12official官方榜source ↗×2
16Qwen3.7 MaxQwen1474442026-08-12official官方榜source ↗×6
17Claude Opus 4.8Anthropic1474⚡high44official官方榜source ↗×5⚠±38
18GLM 5.2Z.ai1471432026-08-12official官方榜source ↗×3
19Grok 4.5xAI1469432026-08-12official官方榜source ↗×4⚠±50
20MiMo-V2.5-ProXiaomi1468422026-08-12official官方榜source ↗
21GLM 5.1Z.ai1467422026-08-12official官方榜source ↗×4
22GPT-5.6 Terra ProOpenAI1464⚡max412026-08-12official官方榜source ↗×2⚠±58
23Claude Sonnet 5Anthropic1462⚡high412026-08-12official官方榜source ↗×6
24Kimi K2.7 CodeMoonshot AI146141official官方榜source ↗
25DeepSeek V4 Pro 0423DeepSeek1458402026-08-12official官方榜source ↗×5⚠±39.95
26Qwen3.7 PlusQwen1458402026-08-12official官方榜source ↗×3
27Gemini 3.5 Flash LiteGoogle1458402026-08-12official官方榜source ↗×3
28Hy3Tencent1457402026-08-12official官方榜source ↗
29GPT-5.6 LunaOpenAI145238official官方榜source ↗
30GPT-5.6 Luna ProOpenAI1450⚡max382026-08-12official官方榜source ↗
31MiniMax M3MiniMax1444362026-08-12official官方榜source ↗×4
32Grok 4.3xAI1442362026-08-12official官方榜source ↗×3⚠±95
33DeepSeek V4 Flash 0423DeepSeek1435342026-08-12official官方榜source ↗×3⚠±31
34Gemini 3.1 Flash LiteGoogle1432332026-08-12official官方榜source ↗×4⚠±41
35Kimi K2 ThinkingMoonshot AI1430⚡high332026-08-12official官方榜source ↗×2
36Nemotron 3 UltraNVIDIA1427322026-08-12official官方榜source ↗
37Mistral Medium 3.5Mistral AI1427322026-08-12official官方榜source ↗×2
38DeepSeek V3.2DeepSeek1425322026-08-12official官方榜source ↗×4⚠±43
39Llama 4 MaverickMeta141730official官方榜source ↗×9
40MiniMax M2.7MiniMax1416292026-08-12official官方榜source ↗×2
41Mistral LargeMistral AI1415292026-08-12official官方榜source ↗×3⚠±60
42Claude Haiku 4.5Anthropic1413282026-08-12official官方榜source ↗×5⚠±90
43Step 3.7 FlashStepFun139524official官方榜source ↗×2
44Llama 4 ScoutMeta1390233rd-party第三方source ↗×2⚠±109
45gpt-oss-120bOpenAI1365173rd-party第三方source ↗
46Command ACohere135414official官方榜source ↗×2
47Seed-2.0-LiteByteDance Seed135213official官方榜source ↗
48Phi 4Microsoft125603rd-party第三方source ↗×2

How scores become the Intelligence Score →成绩如何合成智能评分 →