Head-to-head正面对比
Grok 4.5 vs Gemini 3.1 Pro Preview
xAI meets Google: sourced benchmark scores side by side, dimension by dimension, with prices and specs. Every conclusion below is derived from the published dataset — no opinions, no sponsorship.
xAI 对 Google:带来源的 benchmark 成绩逐项并排,附价格与规格。下面每条结论都从公开数据集推导——没有主观意见,没有厂商赞助。
◈
The verdict, derived from data结论(全部由数据推导)
- On overall score they are nearly tied (80.5 vs 82.4) — decide on price, context and workload mix.总分几乎打平(80.5 对 82.4)——按价格、上下文和任务结构来选。
- Grok 4.5 is stronger in math.Grok 4.5 在数学上更强。
- Gemini 3.1 Pro Preview is stronger in reasoning, agent, preference.Gemini 3.1 Pro Preview 在推理、智能体、偏好上更强。
- Grok 4.5 is 1.8× cheaper on blended price ($4.00 vs $7.00 per 1M tokens, in/out average).混合价(输入输出均值)上 Grok 4.5 便宜 1.8 倍:$4.00 对 $7.00/百万 token。
- Budget pick: Grok 4.5 delivers 98% of the score at 57% of the blended cost.预算优先选 Grok 4.5:用 57% 的混合成本拿到 98% 的分数。
- Long documents: Gemini 3.1 Pro Preview has the much larger context window (1M vs 500K tokens).长文档选 Gemini 3.1 Pro Preview:上下文窗口大得多(1M 对 500K token)。
● Grok 4.5 ● Gemini 3.1 Pro Preview · missing dimensions count as 0 in the chart缺数据的维度在图上按 0 计
≡
Spec by spec逐项规格
| Grok 4.5 | Gemini 3.1 Pro Preview | |
|---|---|---|
| Provider厂商 | xAI | |
| Released发布 | 2026-07-08 | 2026-02-19 |
| Intelligence Score智能评分 | 80.5 | 82.4 |
| Benchmark coveragebenchmark 覆盖 | 95% | 100% |
| Reasoning推理 | 83.7 | 90.5 |
| Coding代码 | 88.6 | 85.6 |
| Knowledge知识 | 93.9 | 94.5 |
| Math数学 | 74.1 | 58.5 |
| Agent智能体 | 38 | 85.7 |
| Preference偏好 | 42.5 | 46.8 |
| API input $/1M输入价 $/百万 | $2.00 | $2.00 |
| API output $/1M输出价 $/百万 | $6.00 | $12 |
| Blended $/1M混合价 $/百万 | $4.00 | $7.00 |
| Value (score per $)性价比(分数/美元) | 20.1 | 11.8 |
| Context window上下文 | 500K | 1M |
| Max output最大输出 | 450K | 66K |
| Reasoning tiers推理档位 | low / medium / high (cannot disable) | minimal / low / medium / high |
| Open weights开放权重 | No否 | No否 |
| CN-direct国内直连 | No否 | No否 |
| Free tier免费层 | No否 | No否 |
⚔
Benchmark by benchmark逐项 benchmark
| Benchmark基准 | Grok 4.5 | Gemini 3.1 Pro Preview |
|---|---|---|
| Humanity's Last Exam | 40.3 | 46.44⚡high |
| GPQA Diamond | 92.9 | 95.5 |
| MMLU-Pro | 89.2 | 89.8 |
| SWE-bench Verified | 86.6 | 80.6 |
| AIME 2025 | 91.7⚡high | 93.3 |
| LiveCodeBench | 79 | 91.7 |
| FrontierMath | 48 | 16.7 |
| Terminal-Bench | 83.3 | 68.5 |
| LMArena (Chatbot Arena) | 1469 | 1486 |
| WebArena | 35 | 69 |
Grok 4.5 dossier → 档案 → · Gemini 3.1 Pro Preview dossier → 档案 → · How we score评分方法