Head-to-head正面对比
Llama 4 Maverick vs Qwen3.7 Plus
Meta meets Qwen: sourced benchmark scores side by side, dimension by dimension, with prices and specs. Every conclusion below is derived from the published dataset — no opinions, no sponsorship.
Meta 对 Qwen:带来源的 benchmark 成绩逐项并排,附价格与规格。下面每条结论都从公开数据集推导——没有主观意见,没有厂商赞助。
◈
The verdict, derived from data结论(全部由数据推导)
- Qwen3.7 Plus leads overall: 81.7 vs 41.1 on the Intelligence Score.总分上 Qwen3.7 Plus 领先:81.7 对 41.1。
- Qwen3.7 Plus is stronger in reasoning, coding, knowledge, math, agent, preference.Qwen3.7 Plus 在推理、代码、知识、数学、智能体、偏好上更强。
- Llama 4 Maverick is 1.8× cheaper on blended price ($0.448 vs $0.800 per 1M tokens, in/out average).混合价(输入输出均值)上 Llama 4 Maverick 便宜 1.8 倍:$0.448 对 $0.800/百万 token。
- Budget pick: Llama 4 Maverick delivers 50% of the score at 56% of the blended cost.预算优先选 Llama 4 Maverick:用 56% 的混合成本拿到 50% 的分数。
- Need CN-direct endpoints: pick Qwen3.7 Plus.需要国内直连:选 Qwen3.7 Plus。
● Llama 4 Maverick ● Qwen3.7 Plus · missing dimensions count as 0 in the chart缺数据的维度在图上按 0 计
≡
Spec by spec逐项规格
| Llama 4 Maverick | Qwen3.7 Plus | |
|---|---|---|
| Provider厂商 | Meta | Qwen |
| Released发布 | 2025-04-05 | 2026-06-03 |
| Intelligence Score智能评分 | 41.1 | 81.7 |
| Benchmark coveragebenchmark 覆盖 | 95% | 76% |
| Reasoning推理 | 41.5 | 77.5 |
| Coding代码 | 39.6 | 86.9 |
| Knowledge知识 | 84.7 | 93.2 |
| Math数学 | 11 | 85.7 |
| Agent智能体 | 18 | 93.8 |
| Preference偏好 | 29.5 | 39.8 |
| API input $/1M输入价 $/百万 | $0.200 | $0.320 |
| API output $/1M输出价 $/百万 | $0.696 | $1.28 |
| Blended $/1M混合价 $/百万 | $0.448 | $0.800 |
| Value (score per $)性价比(分数/美元) | 91.7 | 102.1 |
| Context window上下文 | 1M | 1M |
| Max output最大输出 | 115K | 131K |
| Reasoning tiers推理档位 | — | thinking on / off |
| Open weights开放权重 | Yes支持 | Yes支持 |
| CN-direct国内直连 | No否 | Yes支持 |
| Free tier免费层 | No否 | No否 |
⚔
Benchmark by benchmark逐项 benchmark
| Benchmark基准 | Llama 4 Maverick | Qwen3.7 Plus |
|---|---|---|
| Humanity's Last Exam | 5.68 | 34.7 |
| GPQA Diamond | 69.8 | 90.3 |
| MMLU-Pro | 80.5 | 88.5 |
| AIME 2025 | 20 | 85.7 |
| LiveCodeBench | 43.4 | 89.6 |
| Terminal-Bench | 8.8 | 70.3 |
| τ²-bench | 17.8 | 93 |
| LMArena (Chatbot Arena) | 1417 | 1458 |
Llama 4 Maverick dossier → 档案 → · Qwen3.7 Plus dossier → 档案 → · How we score评分方法