Benchmark Heaven
Price & provider filters · adjusted costs
Global
Applies to price views & model offers; benchmark evidence stays unfiltered

About & data sources

Benchmark Heaven brings together pricing and benchmark data for open-source and frontier LLMs into one comparable view. Prices are normalized to USD per 1M tokens (input and output) unless a platform prices differently (GitHub Copilot's current token/AI-Credit rates and legacy request billing are shown on a separate product axis).

Sources

What are “Featured” models?

★ Featured marks the specific models this tool was commissioned to track closely — the current frontier and leading open-weight families that matter most for the price/capability comparison. Filtering to “Featured” (the default on most pages) hides the long tail of older or niche models so the charts and tables stay focused. The featured set is:

Turn the “Featured” filter off on any page to explore all 835 tracked models. Featured status is derived from the model's family, so every reasoning variant of a featured family (e.g. each GPT-5.5 effort level) is included.

Snapshot

Dataset generated: Fri, 11 Sep 2026 02:14:45 GMT
Models: 835 · families: 650 · offers: 2786 · providers: 89
Per-source collection dates:
  • openrouter: 2026-09-10
  • artificialanalysis: 2026-09-10
  • designarena: 2026-09-10
  • aws_bedrock: 2026-09-08
  • azure_foundry: 2026-09-08
  • google_vertex: 2026-09-08
  • nebius: 2026-09-08
  • inceptron: 2026-09-08
  • scaleway: 2026-09-08
  • ionos: 2026-09-08
  • mistral: 2026-09-08
  • tensorx: 2026-09-08
  • chutes: 2026-09-08
  • ovhcloud: 2026-09-08
  • stackit: 2026-09-08
  • t_systems_llm_hub: 2026-09-08
  • aa_coding_agents: 2026-09-09
  • github_copilot: 2026-09-08
  • claude_code: 2026-09-08
  • aa_coding_agents_v1_5: 2026-09-10
  • provider_meta: 2026-07-12
  • aa_efficiency: 2026-09-10T20:13:54.198Z
  • openrouter_efficiency: 2026-09-10T20:13:54.306Z
  • chutes_efficiency: 2026-09-10T19:44:35.632Z

Methodology

Costs default to modeled USD/task, using exact-variant AA output tokens, OpenRouter usage I/O (Chutes global fallback), and exact-endpoint cache-hit rates. Missing values are assumed and flagged: 1,000 output tokens/task, no cache discount, and zero additional cache writes. General traffic and benchmark tasks are proxies for coding-agent usage. Click an underlined price to inspect the full formula, inputs, dates and assumptions. Raw list-price mode uses your chosen fixed blend per million tokens (including 1:1, 3:1, 10:1, 100:1, input-only and output-only). Capability scores are shown as published; AA indices are 0–100, DesignArena values are Elo. The Composite score uses five slots: AA Coding, source-matched AA Coding Agent, AA Intelligence, DesignArena Frontend and DesignArena Full-Stack. AA values are clamped to 0–100. A DesignArena board qualifies at an app-selected minimum of 200 battles, aligned with the source's typical preliminary/reliability threshold; its Elo is converted to the expected score against a fixed Elo 1000 opponent. Each observed slot is converted to its percentile among the current catalog's unique observed values. Every missing slot inherits that model's mean observed percentile, producing a base score exactly equal to the mean of its available percentiles. A final dominance-safe projection prevents missing data from reversing an otherwise unambiguous comparison: when one model covers every reliable slot of another measured model and is no worse in any shared slot, the catalog scores are adjusted by the smallest symmetric amount needed to keep the dominating model at least 0.1 points ahead. The unadjusted base and any adjustment are shown separately in model details. A model with no reliable observed slot receives the neutral fallback 50; its zero evidence coverage remains distinct from a measured score and is excluded from capability charts. Coding Agent results are attached to an exact or explicitly audited model/reasoning identity; every harness result is retained and their median is used. Family-scoped Intelligence.ai / DesignArena results are attached exactly once to the deterministic collapsed-family representative rather than copied to effort siblings. The provenance note explicitly states that this does not identify the tested effort setting. Raw DesignArena score views continue to show Elo. Stable model ids and repositories are preferred over fuzzy names so distinct releases, modes, context tiers and serving routes do not share the wrong price. See the repository README and data/SCRAPING.md for how each source is collected and refreshed.