The AI tools index that doesn't waste your time.不浪费你时间的 AI 工具索引。
LynxModel

Benchmark explorerBenchmark 百科

The 11 public benchmarks behind the Intelligence Score — what each one actually tests, why it made the cut, who maintains it, and the live standings from our 62 tracked models.

智能评分背后的 11 个公开 benchmark——每个考什么、为什么入选、谁在维护,以及我们追踪的 62 个模型在上面的实时排名。

Reasoning推理 · weight 15%权重 15%

GPQA Diamond

Graduate-level multiple-choice science questions (biology, physics, chemistry) written so that experts with web access still find them hard. 'Diamond' is the highest-quality 198-question subset.

研究生级别的科学选择题(生物/物理/化学),专门设计成'专家带着搜索引擎也觉得难'。Diamond 是其中质量最高的 198 题子集。

60 models scored 个模型有成绩
Reasoning推理 · weight 15%权重 15%

Humanity's Last Exam

3,000 expert-written questions across 100+ disciplines, explicitly designed as the hardest closed-book exam for AI. Frontier models still score well below human experts.

3000 道由各领域专家撰写的闭卷考题,覆盖 100+ 学科,设计目标就是'AI 最难的一张卷子'。顶尖模型得分仍远低于人类专家。

55 models scored 个模型有成绩
Coding代码 · weight 12%权重 12%

SWE-bench Verified

Real GitHub issues from popular Python repos. The model must produce a patch that makes failing tests pass. 'Verified' is the 500-issue human-validated subset.

来自热门 Python 仓库的真实 GitHub issue,模型要交出能让失败测试通过的补丁。Verified 是经过人工校验的 500 题子集。

40 models scored 个模型有成绩
Coding代码 · weight 7%权重 7%

LiveCodeBench

Competitive-programming problems (LeetCode/AtCoder/Codeforces) collected continuously after model training cutoffs, scored by pass rate with zero contamination by design.

持续从 LeetCode/AtCoder/Codeforces 采集的新竞赛题,题目发布晚于模型训练截止日,从机制上杜绝刷题污染,按通过率计分。

44 models scored 个模型有成绩
Coding代码 · weight 6%权重 6%

Terminal-Bench

Real command-line tasks inside Docker containers: compiling, debugging, configuring systems. The model acts as an agent in an actual terminal.

在真实 Docker 容器里完成命令行任务:编译、调试、配环境。模型以 agent 身份在真终端里操作。

40 models scored 个模型有成绩
Knowledge知识 · weight 15%权重 15%

MMLU-Pro

A harder re-build of the classic MMLU: 12K questions, 10 answer choices instead of 4, more reasoning-heavy, 14 domains. The original MMLU is saturated (>93%) and retired from our scoring.

经典 MMLU 的加难重制版:1.2 万题、10 个选项(原来 4 个)、更重推理、14 个学科。原版 MMLU 已饱和(>93%),我们不再计入评分。

30 models scored 个模型有成绩 · ⚠ saturation watch饱和观察
Math数学 · weight 8%权重 8%

AIME 2025

Problems from the 2025 American Invitational Mathematics Examination — olympiad-level math that requires multi-step derivation, not pattern matching.

2025 年美国数学邀请赛真题——奥赛级数学题,需要多步推导而不是套题型。

55 models scored 个模型有成绩 · ⚠ saturation watch饱和观察
Math数学 · weight 7%权重 7%

FrontierMath

Research-level mathematics problems written with active mathematicians, spanning tiers from hard graduate problems to previously unsolved ones.

与在职数学家合作编写的研究级数学题,难度从研究生难题一路到此前未解的问题。

24 models scored 个模型有成绩
Agent智能体 · weight 5%权重 5%

τ²-bench

Tool-calling agent evaluation in realistic customer-service domains (airline, retail, telecom). The agent must use APIs, follow policies, and handle multi-turn user interactions.

仿真客服场景的工具调用评测(航空/零售/电信):agent 必须正确调 API、遵守业务规则、处理多轮用户对话。

30 models scored 个模型有成绩
Agent智能体 · weight 5%权重 5%

WebArena

A self-hosted replica of real websites (shopping, forums, maps, CMS). Agents must complete multi-step tasks by actually clicking and typing in a browser.

自托管的仿真网站环境(购物/论坛/地图/CMS),agent 要在浏览器里真点真输入,完成多步任务。

11 models scored 个模型有成绩
Preference偏好 · weight 5%权重 5%

LMArena (Chatbot Arena)

Millions of blind side-by-side human votes: users compare two anonymous models and pick the better answer. Scores are Bradley-Terry ratings (Elo-like).

数百万次真人盲测投票:用户并排看两个匿名模型的回答,选更好的那个。分数是 Bradley-Terry 评分(类 Elo)。

48 models scored 个模型有成绩