Head-to-head正面对比
Mistral Large vs Command A
Mistral AI meets Cohere: sourced benchmark scores side by side, dimension by dimension, with prices and specs. Every conclusion below is derived from the published dataset — no opinions, no sponsorship.
Mistral AI 对 Cohere:带来源的 benchmark 成绩逐项并排,附价格与规格。下面每条结论都从公开数据集推导——没有主观意见,没有厂商赞助。
◈
The verdict, derived from data结论(全部由数据推导)
- On overall score they are nearly tied (48.7 vs 49.2) — decide on price, context and workload mix.总分几乎打平(48.7 对 49.2)——按价格、上下文和任务结构来选。
- Mistral Large is stronger in coding, preference.Mistral Large 在代码、偏好上更强。
- Command A is stronger in reasoning, math.Command A 在推理、数学上更强。
- Mistral Large is 1.6× cheaper on blended price ($4.00 vs $6.25 per 1M tokens, in/out average).混合价(输入输出均值)上 Mistral Large 便宜 1.6 倍:$4.00 对 $6.25/百万 token。
- Budget pick: Mistral Large delivers 99% of the score at 64% of the blended cost.预算优先选 Mistral Large:用 64% 的混合成本拿到 99% 的分数。
- Long documents: Command A has the much larger context window (256K vs 128K tokens).长文档选 Command A:上下文窗口大得多(256K 对 128K token)。
- Need self-hosting or fine-tuning: Mistral Large ships open weights.需要自托管或微调:Mistral Large 开放权重。
● Mistral Large ● Command A · missing dimensions count as 0 in the chart缺数据的维度在图上按 0 计
≡
Spec by spec逐项规格
| Mistral Large | Command A | |
|---|---|---|
| Provider厂商 | Mistral AI | Cohere |
| Released发布 | 2024-02-26 | 2025-03-13 |
| Intelligence Score智能评分 | 48.7 | 49.2 |
| Benchmark coveragebenchmark 覆盖 | 90% | 82% |
| Reasoning推理 | 26.6 | 36.5 |
| Coding代码 | 52.8 | 31.4 |
| Knowledge知识 | 77 | 74.9 |
| Math数学 | 64.4 | 90 |
| Agent智能体 | — | 85.8 |
| Preference偏好 | 29 | 13.7 |
| API input $/1M输入价 $/百万 | $2.00 | $2.50 |
| API output $/1M输出价 $/百万 | $6.00 | $10 |
| Blended $/1M混合价 $/百万 | $4.00 | $6.25 |
| Value (score per $)性价比(分数/美元) | 12.2 | 7.9 |
| Context window上下文 | 128K | 256K |
| Max output最大输出 | 102K | 8K |
| Reasoning tiers推理档位 | — | — |
| Open weights开放权重 | Yes支持 | No否 |
| CN-direct国内直连 | No否 | No否 |
| Free tier免费层 | No否 | No否 |
⚔
Benchmark by benchmark逐项 benchmark
| Benchmark基准 | Mistral Large | Command A |
|---|---|---|
| Humanity's Last Exam | 4.1 | 11.4 |
| GPQA Diamond | 43.9 | 50.8 |
| MMLU-Pro | 73.11 | 71.2 |
| SWE-bench Verified | 49.5 | 26.8 |
| AIME 2025 | 85 | 90 |
| LiveCodeBench | 36.2 | 35.07 |
| LMArena (Chatbot Arena) | 1415 | 1354 |
Mistral Large dossier → 档案 → · Command A dossier → 档案 → · How we score评分方法