Benchmark Heaven / Model intelligence
Benchmarks in perspective.
Costs in context.
Find the model that meets your benchmark threshold, then compare estimated task costs across providers. Adjusted prices are on by default. Explore versioned results, check their sources, and see where evidence is missing. Raw list prices and fixed input/output blends remain selectable.
Adjusted costs are modeled USD per task: AA output tokens × OpenRouter usage I/O (Chutes global fallback). These are general usage and benchmark proxies for coding-agent work. Missing AA data assumes 1,000 output tokens/task; unknown cache hit assumes 0%; unmeasured additional cache writes assume 0 tokens. Click any underlined price for exact inputs, dates and assumptions. Raw list prices use the selected fixed input/output blend, in USD per million tokens.
Models without AA task-token measurements are excluded from this ranking. Turn off “Measured task tokens only” to include their assumed task costs.
| Top provider channels | |||||
|---|---|---|---|---|---|
| ▸GLM-5.3-Flashopen★ | Z.ai | 77.4 | 25 | DeepInfra/OpenRouter Relace/OpenRouter Wafer/OpenRouter | |
| ▸MiMo-V2.5-Proopen★ | Xiaomi | 60.1 | 5 | GMICloud/OpenRouter DeepInfra/OpenRouter DigitalOcean/OpenRouter | |
| ▸GPT-5.6 Luna★ | OpenAI | 70.8 | 6 | Azure/OpenRouter OpenAI/OpenRouter Azure AI Foundry | |
| ▸Kimi K2.7 Codeopen★ | Moonshot AI | 60.2 | 14 | Inceptron DeepInfra/OpenRouter Inceptron/OpenRouter | |
| ▸GPT-5.6 Sol★ | OpenAI | 77.5 | 6 | OpenAI/OpenRouter Amazon Bedrock/OpenRouter Azure/OpenRouter | |
| ▸GPT-5.6 Terra★ | OpenAI | 65.4 | 6 | Azure AI Foundry Azure/OpenRouter OpenAI/OpenRouter | |
| ▸MiniMax-M3open★ | MiniMax | 66.6 | 13 | CoreWeave/OpenRouter GMICloud/OpenRouter DeepInfra/OpenRouter | |
| ▸GLM-5.1opendeprecated★ | Z.ai | 56.5 | 13 | Chutes Chutes/OpenRouter DeepInfra/OpenRouter | |
| ▸GLM-5.3open★ | Z.ai | 87.1 | 26 | Novita/OpenRouter GMICloud/OpenRouter Phala/OpenRouter | |
| ▸Grok 4.5★ | xAI | 82.8 | 1 | xAI/OpenRouter | |
| ▸DeepSeek V4 Pro 0813open★ | DeepSeek | 65.4 | 16 | Novita/OpenRouter GMICloud/OpenRouter Ionstream/OpenRouter | |
| ▸Muse Spark 1.2★ | Meta | 78.7 | 1 | Meta/OpenRouter | |
| ▸Grok 4.6★ | xAI | 91.2 | 3 | xAI/OpenRouter AWS Bedrock Amazon Bedrock/OpenRouter | |
| ▸Muse Spark 1.3★ | Meta | 95.4 | 1 | Meta/OpenRouter | |
| ▸Kimi K3open★ | Moonshot AI | 92.4 | 18 | Sail Research/OpenRouter Together/OpenRouter Relace/OpenRouter | |
| ▸GPT-5.5deprecated★ | OpenAI | 65.5 | 6 | Azure AI Foundry Azure/OpenRouter OpenAI/OpenRouter | |
| ▸GPT-6 Astra★ | OpenAI | 96.0 | 3 | OpenAI/OpenRouter Azure/OpenRouter | |
| ▸Claude Opus 5★ | Anthropic | 95.7 | 10 | Anthropic/OpenRouter Amazon Bedrock/OpenRouter Google/OpenRouter | |
| ▸Claude Opus 4.8deprecated★ | Anthropic | 84.3 | 9 | Google Vertex AI Anthropic Claude Platform on AWS/OpenRouter | |
| ▸Claude Sonnet 5★ | Anthropic | 85.8 | 9 | Google Vertex AI Anthropic Claude Platform on AWS/OpenRouter | |
| ▸Claude Fable 5.1★ | Anthropic | 100.0 | 7 | Amazon Bedrock/OpenRouter Anthropic/OpenRouter Google/OpenRouter | |
| ▸Claude Fable 5★ | Anthropic | 93.7 | 8 | Google Vertex AI Anthropic Claude Platform on AWS/OpenRouter |