The AI tools index that doesn't waste your time.不浪费你时间的 AI 工具索引。
Coding · weight 6%代码 · 权重 6%

Terminal-Bench

official site官网

What it tests: Real command-line tasks inside Docker containers: compiling, debugging, configuring systems. The model acts as an agent in an actual terminal.

考什么:在真实 Docker 容器里完成命令行任务:编译、调试、配环境。模型以 agent 身份在真终端里操作。

Why it counts: The fastest-rising agentic-coding benchmark of 2026 — closest proxy for 'can this model drive my dev environment'.

为什么算数:2026 年上升最快的代码 benchmark——最接近'这个模型能不能直接开动我的开发环境'的真实体验。

Maintainer维护方Harbor Framework 团队
License许可MIT
Reference point (prior-gen SOTA)参考点(上一代 SOTA)60

Standings实时排名

40 models with sourced scores 个模型有溯源成绩
#Model模型Score成绩Index指数Dated日期Measured by测评方
1GPT-5.6 SolOpenAI91.9100official官方榜source ↗×9⚠±57.3
2Kimi K3Moonshot AI88.396official官方榜source ↗×6
3Qwen3.8 Max (0902)Qwen86.6943rd-party第三方source ↗×3
4Claude Fable 5Anthropic83.891official官方榜source ↗×6⚠±53.9
5Grok 4.5xAI83.391official官方榜source ↗×3
6Muse Spark 1.2meta82.990vendor-reported厂商自报source ↗
7DeepSeek V4 Flash 0731DeepSeek82.7903rd-party第三方source ↗×3
8GPT-5.5OpenAI82.790official官方榜source ↗×4
9DeepSeek V4 Flash 0423DeepSeek82.7902026-07-313rd-party第三方source ↗×3⚠±25.8
10GLM 5.2Z.ai81883rd-party第三方source ↗×5
11Claude Sonnet 5Anthropic80.487official官方榜source ↗×6
12GPT-5.6 TerraOpenAI78.485official官方榜source ↗×6⚠±9.6
13Gemini 3.6 FlashGoogle78852026-073rd-party第三方source ↗×3⚠±4.22
14Gemini 3.5 FlashGoogle76.2832026-073rd-party第三方source ↗×2
15GPT-5.6 LunaOpenAI75.782official官方榜source ↗×6⚠±9
16Claude Opus 4.8Anthropic74.681official官方榜source ↗×5⚠±10.4
17Qwen3.7 PlusQwen70.3773rd-party第三方source ↗
18Qwen3.7 MaxQwen69.7762026-063rd-party第三方source ↗×4⚠±4.8
19Gemini 3.1 Pro PreviewGoogle68.575official官方榜source ↗×6⚠±41.4
20DeepSeek V4 Pro 0423DeepSeek67.9⚡max743rd-party第三方source ↗×5⚠±36.4
21Kimi K2.7 CodeMoonshot AI66.7733rd-party第三方source ↗
22MiniMax M3MiniMax6672official官方榜source ↗×4
23Mistral Medium 3.5Mistral AI66723rd-party第三方source ↗
24Mistral LargeMistral AI66722026-08-013rd-party第三方source ↗×2⚠±42.25
25GLM 5.1Z.ai63.5692026-073rd-party第三方source ↗×4⚠±7.3
26Step 3.7 FlashStepFun59.5653rd-party第三方source ↗
27MiniMax M2.7MiniMax5762official官方榜source ↗×3⚠±17.6
28Gemini 3.5 Flash LiteGoogle54593rd-party第三方source ↗×2
29Grok 4.20xAI47.1⚡high513rd-party第三方source ↗
30DeepSeek V3.2DeepSeek46.450official官方榜source ↗×2
31Seed-2.0-LiteByteDance Seed45493rd-party第三方source ↗
32Claude Opus 5Anthropic43.5⚡max47official官方榜source ↗×7⚠±46.4
33Grok 4.3xAI37.9413rd-party第三方source ↗×2⚠±4.05
34Kimi K2 ThinkingMoonshot AI35.7⚡high39official官方榜source ↗×2
35Claude Haiku 4.5Anthropic27.530official官方榜source ↗×3⚠±13.5
36Mercury 2Inception26.5293rd-party第三方source ↗
37gpt-oss-120bOpenAI18.7203rd-party第三方source ↗
38Gemini 3.1 Flash LiteGoogle17.5193rd-party第三方source ↗
39Llama 4 MaverickMeta8.8103rd-party第三方source ↗
40Llama 4 ScoutMeta8.8103rd-party第三方source ↗

How scores become the Intelligence Score →成绩如何合成智能评分 →