Benchmark Heaven
Price & provider filters · adjusted costs
Global
Applies to price views & model offers; benchmark evidence stays unfiltered
← All models

Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)

★ featured
Anthropic · released 2026-09-01 · 8 recorded token offers

Top 5 cheapest providers (Adjusted $/task)

Within the active global provider, residency and confidentiality filters.

#ProviderPlatformRaw input $/1MRaw output $/1MAdjusted $/task
1Amazon BedrockOpenRouter$10.00$50.00
2AnthropicOpenRouter$10.00$50.00
3GoogleOpenRouter$10.00$50.00
4AzureOpenRouter$10.00$50.00
5AWS BedrockAWS Bedrock$10.00$50.00

Benchmarks

Composite99.6
Composite evidence2/5
AA Coding · unversioned snapshot 2026-09-1080.7
AA Coding Agent · v1.4 · 2026-09-09
AA Intelligence · unversioned snapshot 2026-09-1053.2
DesignArena Frontend · unversioned snapshot 2026-09-10
DesignArena Full-Stack · unversioned snapshot 2026-09-10
Output speed (t/s)59

MODEL EVIDENCE

Benchmark sheet

14 / 73 registered benchmark versions covered · 16 observations including existing index snapshots.

Compare this model ↗
Composite retains Coding Agent v1.4 · source 2026-09-09

The five Composite inputs remain unchanged. Its Coding Agent input is the median across complete harness results in the retained v1.4 snapshot from 2026-09-09. Artificial Analysis now publishes v1.5, with different components. Explore current v1.5 separately. Dated snapshot labels on other indices identify unversioned source captures, not a verified semantic version.

Profile signals

1 threshold-crossing signal flagged · 15 eligible benchmark families

How flags are calculated

Heuristic screen, not statistical significance: benchmark families are correlated and source uncertainty is unknown. Peer evidence needs ≥ 20 independently measured matched configurations from ≥ 10 distinct model families. A flag needs a directed population z-score of magnitude ≥ 1.5 and a gap of ≥ 1.5 from the leave-one-benchmark-family-out mean z in the same direction, over ≥ 5 other benchmark families.

  • unusually strong CritPt (AA) snapshot-2026-09-10
    Why
    Observed: 0.31143 fraction
    Peer mean 0.03942 · peer sd 0.0762 · n 521 · families 367
    Directed z 3.57 · baseline z 1.78 · gap 1.789 · profile n 14
    0.31143 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
    Evidence
    Axis: CritPt (AA) · snapshot-2026-09-10 · Published board
    Exact value: 0.311428571428571 fraction
    Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
    Observation id: aa:9b166bf3-42db-4f63-8338-1c4a1244ffe8:critpt

Protocol-compatible measured/vendor divergences

No verified protocol-compatible vendor/measured pair is available for this model; agreement cannot be assessed.

Agentic

Agentic: every observation for this exact model configuration
Benchmark / versionEvaluation groupScore / provenance
AA-Briefcase

Version snapshot-2026-09-10

Published board

Tests multi-week professional knowledge-work projects with linked tasks and large source collections.

1650.01 Elomeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: AA-Briefcase · snapshot-2026-09-10 · Published board
Exact value: 1650.01 Elo
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:9b166bf3-42db-4f63-8338-1c4a1244ffe8:briefcaseBreakdown.overall.elo
GDPval-AA v2

Version 2

Published board

Tests professional knowledge-work deliverables across occupations using AA's Stirrup harness.

1745.29 Elomeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: GDPval-AA v2 · 2 · Published board
Exact value: 1745.29 Elo
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:9b166bf3-42db-4f63-8338-1c4a1244ffe8:gdpval
Harvey LAB-AA

Version snapshot-2026-09-10

Published board

Tests legal-work deliverables across practice areas on Harvey's private task set.

0.9327 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: Harvey LAB-AA · snapshot-2026-09-10 · Published board
Exact value: 0.932696482848459 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:9b166bf3-42db-4f63-8338-1c4a1244ffe8:harveyLab
Terminal-Bench v2.1 (AA)

Version 2.1

Published board

Tests terminal-based work on the 89-task verified refresh using Terminus 2.

0.91011 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: Terminal-Bench v2.1 (AA) · 2.1 · Published board
Exact value: 0.910112359550562 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:9b166bf3-42db-4f63-8338-1c4a1244ffe8:terminalbenchV21
Terminal-Bench v4.0 (AA)

Version 4.0

Published board

Tests terminal-based work on the 66-task release using mini-SWE-agent v2.4.6.

0.55051 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: Terminal-Bench v4.0 (AA) · 4.0 · Published board
Exact value: 0.55050505050505 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:9b166bf3-42db-4f63-8338-1c4a1244ffe8:terminalbenchV40

Coding

Coding: every observation for this exact model configuration
Benchmark / versionEvaluation groupScore / provenance
SciCode (AA subproblems)

Version 1.0.1

Published board

Tests scientific Python programming with scientist-annotated background information.

0.6088 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: SciCode (AA subproblems) · 1.0.1 · Published board
Exact value: 0.608796296296296 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:9b166bf3-42db-4f63-8338-1c4a1244ffe8:scicode
AA Coding Index

Version snapshot-2026-09-10 (unversioned)

Published board

Published AA index from the retained API snapshot. No verified semantic version was supplied.

80.7 pointsmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: AA Coding Index · snapshot-2026-09-10 (unversioned) · Published board
Exact value: 80.7 points
Observed: 2026-09-10 · publication date: not recorded
Observation id: legacy:aa_coding_index:claude-fable-5.1::xhigh
Protocol: inspect · file data/raw/artificialanalysis.json

Knowledge

Knowledge: every observation for this exact model configuration
Benchmark / versionEvaluation groupScore / provenance
Humanity's Last Exam (AA text-only)

Version snapshot-2026-09-10

Published board

Tests expert-level knowledge on AA's text-only Humanity's Last Exam subset.

0.58712 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: Humanity's Last Exam (AA text-only) · snapshot-2026-09-10 · Published board
Exact value: 0.587117701575533 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:9b166bf3-42db-4f63-8338-1c4a1244ffe8:hle
AA-Omniscience Index

Version snapshot-2026-09-10

Published board

Tests factual reliability while rewarding correct answers and penalizing hallucinations.

42.38333 pointsmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: AA-Omniscience Index · snapshot-2026-09-10 · Published board
Exact value: 42.3833333333333 points
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:9b166bf3-42db-4f63-8338-1c4a1244ffe8:omniscience

Long-context

Long-context: every observation for this exact model configuration
Benchmark / versionEvaluation groupScore / provenance
GDP.pdf (AA)

Version snapshot-2026-09-10

Published board

Tests professional reasoning over long PDFs with AA document preparation and grading.

0.262 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: GDP.pdf (AA) · snapshot-2026-09-10 · Published board
Exact value: 0.262 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:9b166bf3-42db-4f63-8338-1c4a1244ffe8:gdpPdfAllPass
AA-LCR v1.1

Version 1.1

Published board

Tests reasoning across multiple long documents with corrected answer keys and grading.

0.83 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: AA-LCR v1.1 · 1.1 · Published board
Exact value: 0.83 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:9b166bf3-42db-4f63-8338-1c4a1244ffe8:lcr

Reasoning

Reasoning: every observation for this exact model configuration
Benchmark / versionEvaluation groupScore / provenance
AA Intelligence Index

Version snapshot-2026-09-10 (unversioned)

Published board

Published AA index from the retained API snapshot. No verified semantic version was supplied.

53.2 pointsmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: AA Intelligence Index · snapshot-2026-09-10 (unversioned) · Published board
Exact value: 53.2 points
Observed: 2026-09-10 · publication date: not recorded
Observation id: legacy:aa_intelligence_index:claude-fable-5.1::xhigh
Protocol: inspect · file data/raw/artificialanalysis.json

Science

Science: every observation for this exact model configuration
Benchmark / versionEvaluation groupScore / provenance
CritPt (AA)

Version snapshot-2026-09-10

Published board

Tests research-level physics reasoning with Python, symbolic and numerical answers.

0.31143 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: CritPt (AA) · snapshot-2026-09-10 · Published board
Exact value: 0.311428571428571 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:9b166bf3-42db-4f63-8338-1c4a1244ffe8:critpt
GPQA Diamond (AA)

Version snapshot-2026-09-10

Published board

Tests graduate-level biology, physics and chemistry knowledge on the Diamond subset.

0.93434 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: GPQA Diamond (AA) · snapshot-2026-09-10 · Published board
Exact value: 0.934343434343434 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:9b166bf3-42db-4f63-8338-1c4a1244ffe8:gpqa

Tool-use

Tool-use: every observation for this exact model configuration
Benchmark / versionEvaluation groupScore / provenance
AutomationBench-AA

Version 1.0.6

Published board

Tests multi-app SaaS workflows through REST tools on a held-out AutomationBench split.

0.57772 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: AutomationBench-AA · 1.0.6 · Published board
Exact value: 0.57772425084896 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:9b166bf3-42db-4f63-8338-1c4a1244ffe8:automationBenchPartialScore
τ³-Banking (AA)

Version 1.0.1

Published board

Tests banking support agents that retrieve policies and change account state through tools.

0.45773 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: τ³-Banking (AA) · 1.0.1 · Published board
Exact value: 0.457731958762887 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:9b166bf3-42db-4f63-8338-1c4a1244ffe8:tauBanking
Missing coverage · 61 benchmark versions

No result does not mean a zero, or that the model was never tested. Collection failures and disputed versions retain their distinct status.

Variants / reasoning settings

VariantCoding · snapshot 2026-09-10Intelligence · snapshot 2026-09-10
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)81.653.4
Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)80.753.2
Claude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback)79.151.2
Claude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback)77.149.1
Claude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback)75.247.0

GitHub Copilot

Current usage-based billing · model token cost is converted to AI Credits at 1 credit = $0.01.

Input $10.00 / 1M
Cached input $0.250 / 1M
Cache write $12.50 / 1M
Output $50.00 / 1M
Status GA

Token offers by platform — Adjusted $/task

Adjusted costs are modeled USD per task: AA output tokens × OpenRouter usage I/O (Chutes global fallback). These are general usage and benchmark proxies for coding-agent work. Missing AA data assumes 1,000 output tokens/task; unknown cache hit assumes 0%; unmeasured additional cache writes assume 0 tokens. Click any underlined price for exact inputs, dates and assumptions. Raw list prices use the selected fixed input/output blend, in USD per million tokens.

8 offers within the active global filters; “—” means the catalog is active but no public token price is available.

OpenRouter (4)

Amazon Bedrockglobalamazon-bedrock$10.00 raw in $/1M$50.00 raw out $/1M
Anthropicglobalanthropic$10.00 raw in $/1M$50.00 raw out $/1M
Googleglobalgoogle-vertex/global$10.00 raw in $/1M$50.00 raw out $/1M
Azureglobalazure$10.00 raw in $/1M$50.00 raw out $/1M

AWS Bedrock (1)

AWS Bedrockglobal$10.00 raw in $/1M$50.00 raw out $/1M

Google Vertex AI (2)

Google Vertex AIglobal$10.00 raw in $/1M$50.00 raw out $/1M
Google Vertex AIeuEU$11.00 raw in $/1M$55.00 raw out $/1M

Anthropic (1)

Anthropicglobal$10.00 raw in $/1M$50.00 raw out $/1M
Benchmark Heaven — Model benchmarks & costs