Model dossier模型档案
Grok 4.20
Grok 4.20 is a high-tier model from xAI (Intelligence Score 80.1/100). Strongest in knowledge (100) and reasoning (91.8); best fit: knowledge QA and long-form writing. It trails in preference (44). Mid-priced: $1.25 in / $2.50 out per 1M tokens. Context window 2M tokens — long-document friendly.
Grok 4.20 是 xAI 的高水平梯队模型(智能评分 80.1/100)。 最强项是知识(100分),其次是推理(91.8分);适合知识问答与长文写作。 短板是偏好(44分)。 定价属中档定价:每百万 token 输入 $1.25 / 输出 $2.50。 上下文 2M token,长文档友好。
80.1
Intelligence Score · benchmark coverage 智能评分 · benchmark 覆盖 90%
◈
Benchmark scoresBenchmark 成绩
9 entries 条| Benchmark基准 | Score成绩 | Index指数 | Dated日期 | Measured by测评方 | |
|---|---|---|---|---|---|
| Humanity's Last ExamReasoning | 50.7 | 88 | — | 3rd-party第三方 | source ↗×2⚠±18.5 |
| GPQA DiamondReasoning | 91 | 95 | — | 3rd-party第三方 | source ↗×4⚠±3.5 |
| MMLU-ProKnowledge | 95 | 100 | — | vendor-reported厂商自报 | source ↗ |
| SWE-bench VerifiedCoding | 76.7⚡high | 80 | 2026-04-11 | 3rd-party第三方 | source ↗×3 |
| AIME 2025Math | 91.7 | 92 | — | 3rd-party第三方 | source ↗×3 |
| LiveCodeBenchCoding | 84.27 | 90 | — | 3rd-party第三方 | source ↗ |
| FrontierMathMath | 14 | 16 | — | 3rd-party第三方 | source ↗ |
| Terminal-BenchCoding | 47.1⚡high | 51 | — | 3rd-party第三方 | source ↗ |
| LMArena (Chatbot Arena)Preference | 1475 | 44 | 2026-08-12 | official官方榜 | source ↗×2 |
≡
Specs & pricing规格与价格
| Provider厂商 | xAI |
|---|---|
| Released发布 | 2026-03-31 |
| Context window上下文 | 2M tokens |
| Max output最大输出 | 1.8M tokens |
| Modality模态 | text+image+file->text |
| Reasoning tiers推理档位 | ⚡low ⚡medium ⚡high default默认 high · low / medium / high (cannot disable) |
| API input price输入价格 | $1.25 / 1M tokens |
| API output price输出价格 | $2.50 / 1M tokens |
⛓
Tools using this model使用它的工具
How we score →评分方法 → · Catalog & pricing via the official OpenRouter API.模型目录与价格来自 OpenRouter 官方 API。