Benchmark explorerBenchmark 百科
The 11 public benchmarks behind the Intelligence Score — what each one actually tests, why it made the cut, who maintains it, and the live standings from our 62 tracked models.
智能评分背后的 11 个公开 benchmark——每个考什么、为什么入选、谁在维护,以及我们追踪的 62 个模型在上面的实时排名。
GPQA Diamond
Graduate-level multiple-choice science questions (biology, physics, chemistry) written so that experts with web access still find them hard. 'Diamond' is the highest-quality 198-question subset.
研究生级别的科学选择题(生物/物理/化学),专门设计成'专家带着搜索引擎也觉得难'。Diamond 是其中质量最高的 198 题子集。
60 models scored 个模型有成绩Humanity's Last Exam
3,000 expert-written questions across 100+ disciplines, explicitly designed as the hardest closed-book exam for AI. Frontier models still score well below human experts.
3000 道由各领域专家撰写的闭卷考题,覆盖 100+ 学科,设计目标就是'AI 最难的一张卷子'。顶尖模型得分仍远低于人类专家。
55 models scored 个模型有成绩SWE-bench Verified
Real GitHub issues from popular Python repos. The model must produce a patch that makes failing tests pass. 'Verified' is the 500-issue human-validated subset.
来自热门 Python 仓库的真实 GitHub issue,模型要交出能让失败测试通过的补丁。Verified 是经过人工校验的 500 题子集。
40 models scored 个模型有成绩LiveCodeBench
Competitive-programming problems (LeetCode/AtCoder/Codeforces) collected continuously after model training cutoffs, scored by pass rate with zero contamination by design.
持续从 LeetCode/AtCoder/Codeforces 采集的新竞赛题,题目发布晚于模型训练截止日,从机制上杜绝刷题污染,按通过率计分。
44 models scored 个模型有成绩Terminal-Bench
Real command-line tasks inside Docker containers: compiling, debugging, configuring systems. The model acts as an agent in an actual terminal.
在真实 Docker 容器里完成命令行任务:编译、调试、配环境。模型以 agent 身份在真终端里操作。
40 models scored 个模型有成绩MMLU-Pro
A harder re-build of the classic MMLU: 12K questions, 10 answer choices instead of 4, more reasoning-heavy, 14 domains. The original MMLU is saturated (>93%) and retired from our scoring.
经典 MMLU 的加难重制版:1.2 万题、10 个选项(原来 4 个)、更重推理、14 个学科。原版 MMLU 已饱和(>93%),我们不再计入评分。
30 models scored 个模型有成绩 · ⚠ saturation watch饱和观察AIME 2025
Problems from the 2025 American Invitational Mathematics Examination — olympiad-level math that requires multi-step derivation, not pattern matching.
2025 年美国数学邀请赛真题——奥赛级数学题,需要多步推导而不是套题型。
55 models scored 个模型有成绩 · ⚠ saturation watch饱和观察FrontierMath
Research-level mathematics problems written with active mathematicians, spanning tiers from hard graduate problems to previously unsolved ones.
与在职数学家合作编写的研究级数学题,难度从研究生难题一路到此前未解的问题。
24 models scored 个模型有成绩τ²-bench
Tool-calling agent evaluation in realistic customer-service domains (airline, retail, telecom). The agent must use APIs, follow policies, and handle multi-turn user interactions.
仿真客服场景的工具调用评测(航空/零售/电信):agent 必须正确调 API、遵守业务规则、处理多轮用户对话。
30 models scored 个模型有成绩WebArena
A self-hosted replica of real websites (shopping, forums, maps, CMS). Agents must complete multi-step tasks by actually clicking and typing in a browser.
自托管的仿真网站环境(购物/论坛/地图/CMS),agent 要在浏览器里真点真输入,完成多步任务。
11 models scored 个模型有成绩LMArena (Chatbot Arena)
Millions of blind side-by-side human votes: users compare two anonymous models and pick the better answer. Scores are Bradley-Terry ratings (Elo-like).
数百万次真人盲测投票:用户并排看两个匿名模型的回答,选更好的那个。分数是 Bradley-Terry 评分(类 Elo)。
48 models scored 个模型有成绩