Head-to-head正面对比
Claude Sonnet 5 vs GPT-5.5
Anthropic meets OpenAI: sourced benchmark scores side by side, dimension by dimension, with prices and specs. Every conclusion below is derived from the published dataset — no opinions, no sponsorship.
Anthropic 对 OpenAI:带来源的 benchmark 成绩逐项并排,附价格与规格。下面每条结论都从公开数据集推导——没有主观意见,没有厂商赞助。
◈
The verdict, derived from data结论(全部由数据推导)
- On overall score they are nearly tied (83.6 vs 82.9) — decide on price, context and workload mix.总分几乎打平(83.6 对 82.9)——按价格、上下文和任务结构来选。
- Claude Sonnet 5 is stronger in reasoning, agent.Claude Sonnet 5 在推理、智能体上更强。
- GPT-5.5 is stronger in math, preference.GPT-5.5 在数学、偏好上更强。
- Claude Sonnet 5 is 2.9× cheaper on blended price ($6.00 vs $18 per 1M tokens, in/out average).混合价(输入输出均值)上 Claude Sonnet 5 便宜 2.9 倍:$6.00 对 $18/百万 token。
● Claude Sonnet 5 ● GPT-5.5 · missing dimensions count as 0 in the chart缺数据的维度在图上按 0 计
≡
Spec by spec逐项规格
| Claude Sonnet 5 | GPT-5.5 | |
|---|---|---|
| Provider厂商 | Anthropic | OpenAI |
| Released发布 | 2026-06-30 | 2026-04-24 |
| Intelligence Score智能评分 | 83.6 | 82.9 |
| Benchmark coveragebenchmark 覆盖 | 100% | 100% |
| Reasoning推理 | 89.1 | 80.5 |
| Coding代码 | 88.2 | 90.6 |
| Knowledge知识 | 92.2 | 90.6 |
| Math数学 | 76.2 | 97.9 |
| Agent智能体 | 75.5 | 55.4 |
| Preference偏好 | 40.8 | 45.8 |
| API input $/1M输入价 $/百万 | $2.00 | $5.00 |
| API output $/1M输出价 $/百万 | $10 | $30 |
| Blended $/1M混合价 $/百万 | $6.00 | $18 |
| Value (score per $)性价比(分数/美元) | 13.9 | 4.7 |
| Context window上下文 | 1M | 1.1M |
| Max output最大输出 | 128K | 128K |
| Reasoning tiers推理档位 | low / medium / high | minimal / low / medium / high / xhigh |
| Open weights开放权重 | No否 | No否 |
| CN-direct国内直连 | No否 | No否 |
| Free tier免费层 | No否 | No否 |
⚔
Benchmark by benchmark逐项 benchmark
| Benchmark基准 | Claude Sonnet 5 | GPT-5.5 |
|---|---|---|
| Humanity's Last Exam | 57.4 | 36.24 |
| GPQA Diamond | 74.6 | 93.5⚡high |
| MMLU-Pro | 87.55 | 86.1 |
| SWE-bench Verified | 85.2 | 87 |
| AIME 2025 | 57.4 | 100 |
| LiveCodeBench | 82.43 | 85.3 |
| FrontierMath | 87 | 85⚡max |
| Terminal-Bench | 80.4 | 82.7 |
| τ²-bench | 80.4 | 46.39 |
| LMArena (Chatbot Arena) | 1462⚡high | 1482⚡high |
| WebArena | 64.5 | 59 |
Claude Sonnet 5 dossier → 档案 → · GPT-5.5 dossier → 档案 → · How we score评分方法