Charts
Capability leaderboards, cheapest-model rankings, and an open-weights vs closed comparison. Switch the score and toggle the featured set. For the price/capability scatter, see Cost vs Capability.
Capability leaderboard — Composite (coverage-neutral, dominance-safe percentiles, 0–100) · 5 fixed inputs; Coding Agent v1.4
Cheapest models — Adjusted $/task
Plotted model costs (keyboard-accessible table, 18 rows, Adjusted $/task)
| Model | Adjusted $/task |
|---|---|
| Kimi K2.5 | |
| Kimi K2.6 | |
| DeepSeek V4 Pro | |
| GLM-5.2 | |
| GPT-5.4 | |
| Claude Opus 4.7 | |
| Claude Opus 4.6 | |
| GLM-5.3-Flash | |
| MiMo-V2.5-Pro | |
| Claude Sonnet 4.6 | |
| GPT-5.6 Luna | |
| Kimi K2.7 Code | |
| GPT-5.6 Sol | |
| GPT-5.6 Terra | |
| MiniMax-M3 | |
| GLM-5.1 | |
| GLM-5.3 | |
| Grok 4.5 |
Open weights vs closed — average capability
Open weights vs closed — average cost (Adjusted $/task)
Constituent models behind each average — arithmetic mean = sum ÷ count
Open: $7.037 ÷ 12 = $0.58638 USD/task (arithmetic mean of the 12 priced models below)
| GLM-5.3-Flash | |
| DeepSeek V4 Pro 0813 | |
| DeepSeek V4 Pro | |
| MiniMax-M3 | |
| Kimi K2.7 Code | |
| Kimi K3 | |
| MiMo-V2.5-Pro | |
| GLM-5.2 | |
| GLM-5.3 | |
| Kimi K2.6 | |
| Kimi K2.5 | |
| GLM-5.1 |
Closed: $71.25 ÷ 18 = $3.958 USD/task (arithmetic mean of the 18 priced models below)
| GPT-6 Astra | |
| GPT-5.6 Luna | |
| GPT-5.6 Sol | |
| GPT-5.6 Terra | |
| Muse Spark 1.2 | |
| Muse Spark 1.3 | |
| Claude Sonnet 5 | |
| Claude Fable 5 | |
| Claude Opus 5 | |
| Claude Fable 5.1 | |
| Grok 4.6 | |
| Grok 4.5 | |
| GPT-5.4 | |
| GPT-5.5 | |
| Claude Opus 4.8 | |
| Claude Opus 4.7 | |
| Claude Opus 4.6 | |
| Claude Sonnet 4.6 |
Adjusted costs are modeled USD per task: AA output tokens × OpenRouter usage I/O (Chutes global fallback). These are general usage and benchmark proxies for coding-agent work. Missing AA data assumes 1,000 output tokens/task; unknown cache hit assumes 0%; unmeasured additional cache writes assume 0 tokens. Click any underlined price for exact inputs, dates and assumptions. Raw list prices use the selected fixed input/output blend, in USD per million tokens.