Head-to-head正面对比
GPT-5.6 Luna vs Gemini 3.1 Flash Lite
OpenAI meets Google: sourced benchmark scores side by side, dimension by dimension, with prices and specs. Every conclusion below is derived from the published dataset — no opinions, no sponsorship.
OpenAI 对 Google:带来源的 benchmark 成绩逐项并排,附价格与规格。下面每条结论都从公开数据集推导——没有主观意见,没有厂商赞助。
◈
The verdict, derived from data结论(全部由数据推导)
- GPT-5.6 Luna leads overall: 83.3 vs 55.7 on the Intelligence Score.总分上 GPT-5.6 Luna 领先:83.3 对 55.7。
- GPT-5.6 Luna is stronger in reasoning, coding, math, preference.GPT-5.6 Luna 在推理、代码、数学、偏好上更强。
● GPT-5.6 Luna ● Gemini 3.1 Flash Lite · missing dimensions count as 0 in the chart缺数据的维度在图上按 0 计
≡
Spec by spec逐项规格
| GPT-5.6 Luna | Gemini 3.1 Flash Lite | |
|---|---|---|
| Provider厂商 | OpenAI | |
| Released发布 | 2026-07-09 | 2026-05-07 |
| Intelligence Score智能评分 | 83.3 | 55.7 |
| Benchmark coveragebenchmark 覆盖 | 68% | 64% |
| Reasoning推理 | 80.1 | 60.5 |
| Coding代码 | 91.9 | 19 |
| Knowledge知识 | — | 82.1 |
| Math数学 | 94.5 | 30 |
| Agent智能体 | — | — |
| Preference偏好 | 38.2 | 33.2 |
| API input $/1M输入价 $/百万 | $0.200 | $0.250 |
| API output $/1M输出价 $/百万 | $1.20 | $1.50 |
| Blended $/1M混合价 $/百万 | $0.700 | $0.875 |
| Value (score per $)性价比(分数/美元) | 119 | 63.7 |
| Context window上下文 | 1.1M | 1M |
| Max output最大输出 | 128K | 66K |
| Reasoning tiers推理档位 | none / minimal / low / medium / high / xhigh | minimal / low / medium / high |
| Open weights开放权重 | No否 | No否 |
| CN-direct国内直连 | No否 | No否 |
| Free tier免费层 | No否 | No否 |
⚔
Benchmark by benchmark逐项 benchmark
| Benchmark基准 | GPT-5.6 Luna | Gemini 3.1 Flash Lite |
|---|---|---|
| Humanity's Last Exam | 37.2 | 17.2 |
| GPQA Diamond | 91.1⚡max | 86.9 |
| AIME 2025 | 100 | 30 |
| Terminal-Bench | 75.7 | 17.5 |
| LMArena (Chatbot Arena) | 1452 | 1432 |
GPT-5.6 Luna dossier → 档案 → · Gemini 3.1 Flash Lite dossier → 档案 → · How we score评分方法