Benchmark Heaven
Price & provider filters · adjusted costs
Global
Applies to price views & model offers; benchmark evidence stays unfiltered

Cost vs Capability

Every model plotted by price (x) against capability (y). The x-axis is inverted — cheaper to the right — and defaults to the cheapest modeled USD per task. The global toggle restores raw list prices per million tokens using your chosen fixed input/output blend; the y-axis is your chosen benchmark score. The best value sits in the upper-right.

Capability (Y): Composite (coverage-neutral, dominance-safe percentiles, 0–100) · 5 fixed inputs; Coding Agent v1.430 models · X inverted: cheaper → right · provider-filtered
closed lab open weightsAnthropicOpenAIZ.aiMoonshot AIMetaDeepSeekxAIMiniMaxXiaomi

Up & to the right is better: more capability for less money. The green Pareto frontier marks and connects the best-value models — those no other model beats on both price and capability. Click any point to open the model detail.

Adjusted costs are modeled USD per task: AA output tokens × OpenRouter usage I/O (Chutes global fallback). These are general usage and benchmark proxies for coding-agent work. Missing AA data assumes 1,000 output tokens/task; unknown cache hit assumes 0%; unmeasured additional cache writes assume 0 tokens. Click any underlined price for exact inputs, dates and assumptions. Raw list prices use the selected fixed input/output blend, in USD per million tokens.

Model prices and scores (accessible table, 30 rows)
ModelComposite (coverage-neutral, dominance-safe percentiles, 0–100) · 5 fixed inputs; Coding Agent v1.4Adjusted $/task
Kimi K2.5 Moonshot AI · open50.8
Kimi K2.6 Moonshot AI · open61.9
DeepSeek V4 Pro DeepSeek · open60.0
GLM-5.2 Z.ai · open70.9
GPT-5.4 OpenAI65.4
Claude Opus 4.7 Anthropic77.4
Claude Opus 4.6 Anthropic70.8
GLM-5.3-Flash Z.ai · open77.4
MiMo-V2.5-Pro Xiaomi · open60.1
Claude Sonnet 4.6 Anthropic60.3
GPT-5.6 Luna OpenAI70.8
Kimi K2.7 Code Moonshot AI · open60.2
GPT-5.6 Sol OpenAI77.5
GPT-5.6 Terra OpenAI65.4
MiniMax-M3 MiniMax · open66.6
GLM-5.1 Z.ai · open56.5
GLM-5.3 Z.ai · open87.1
Grok 4.5 xAI82.8
DeepSeek V4 Pro 0813 DeepSeek · open65.4
Muse Spark 1.2 Meta78.7
Grok 4.6 xAI91.2
Muse Spark 1.3 Meta95.4
Kimi K3 Moonshot AI · open92.4
GPT-5.5 OpenAI65.5
GPT-6 Astra OpenAI96.0
Claude Opus 5 Anthropic95.7
Claude Opus 4.8 Anthropic84.3
Claude Sonnet 5 Anthropic85.8
Claude Fable 5.1 Anthropic100.0
Claude Fable 5 Anthropic93.7