Independent benchmarks · captured Sep 2, 2026

Fable 5.1 benchmarks

Everything on this page comes from Artificial Analysis and the DeepSWE leaderboard - independent third parties, not Anthropic's own eval numbers. Scores move; the capture date above is when these were read.

The headline

Intelligence Index
66#1 of 192 models in its class, on Intelligence Index v4.1.1.
Cost per index task
$3.69#84 of 192 - and up from $3.14 on Fable 5. See the pricing section below.
Output speed
66.4 tok/s#85 of 192 - slower than the 71 tok/s median.
Verbosity
140M tokensOutput tokens across the whole index, against a 71M median. The most verbose model measured.

Artificial Analysis summarises it as "amongst the leading models in intelligence, but particularly expensive when comparing to other models of similar price... also slower than average and very verbose." Running the full index on Fable 5.1 cost them $8,523.16.

Against the rest of the frontier

0204060Claude Fable 5.1 (max)Claude Fable 5.1 (max): Intelligence Index 6666Claude Fable 5.1 (xhigh)Claude Fable 5.1 (xhigh): Intelligence Index 6565Claude Opus 5 (max)Claude Opus 5 (max): Intelligence Index 6363Claude Fable 5 (max)Claude Fable 5 (max): Intelligence Index 6262GPT-5.6 Sol (max)GPT-5.6 Sol (max): Intelligence Index 6161Grok 4.6 (high)Grok 4.6 (high): Intelligence Index 6161Kimi K3 (max)Kimi K3 (max): Intelligence Index 6060GLM-5.3 (max)GLM-5.3 (max): Intelligence Index 6060
Artificial Analysis Intelligence Index v4.1.1, top models by score. Fable 5.1 takes the top two slots, but the four-point lead over Fable 5 is worth reading next to the cost column in the table below: GPT-5.6 Sol is within five points at roughly a quarter of the cost per task. Source: Artificial Analysis leaderboard.
ModelIntelligence IndexCost per taskMedian tok/sContext
Claude Fable 5.1 (max)66$3.69661M
Claude Fable 5.1 (xhigh)65$2.65591M
Claude Opus 5 (max)63$2.34521M
Claude Fable 5 (max)62$3.14651M
GPT-5.6 Sol (max)61$0.95741M
Grok 4.6 (high)61$0.9453500k
Kimi K3 (max)60$0.84381.05M
GLM-5.3 (max)60$0.68701M

Intelligence by effort level

Fable 5.1 is measured at every effort setting, and the two panels below do not rise at the same rate. Going from low to max buys eight index points and costs 4.8× as much per task.

INTELLIGENCE INDEX03570low: 5858lowmedium: 6060mediumhigh: 6262highxhigh: 6565xhighmax: 6666maxCOST PER INDEX TASK$0$2$4low: $0.77$0.77lowmedium: $1.00$1.00mediumhigh: $1.43$1.43highxhigh: $2.65$2.65xhighmax: $3.69$3.69max
Left: Intelligence Index by effort level. Right: cost per index task, same effort levels, plotted on its own axis because the two measures share no scale. Low effort already scores 58 - within eight points of max - for 21% of the cost, which is the practical argument against defaulting everything to max. All levels are the "with fallback" configurations, meaning safeguard-blocked requests are answered by another Claude model rather than scored zero.

Per-evaluation breakdown

Fable 5.1 leads or ties Fable 5 and Opus 5 on every evaluation in the index. The widest gaps are agentic: 𝜏³-Banking (tool use) and Terminal-Bench v2.1 (terminal coding).

Fable 5.1 Fable 5 Opus 5
0%25%50%75%100%GPQA DiamondFable 5 on GPQA Diamond: 93%Opus 5 on GPQA Diamond: 93%Fable 5.1 on GPQA Diamond: 94%94%Terminal-Bench v2.1Fable 5 on Terminal-Bench v2.1: 85%Opus 5 on Terminal-Bench v2.1: 89%Fable 5.1 on Terminal-Bench v2.1: 91%91%AA-LCRFable 5 on AA-LCR: 77%Opus 5 on AA-LCR: 76%Fable 5.1 on AA-LCR: 80%80%SciCodeFable 5 on SciCode: 60%Opus 5 on SciCode: 56%Fable 5.1 on SciCode: 62%62%Humanity's Last ExamFable 5 on Humanity's Last Exam: 55%Opus 5 on Humanity's Last Exam: 55%Fable 5.1 on Humanity's Last Exam: 59%59%τ³-BankingFable 5 on τ³-Banking: 38%Opus 5 on τ³-Banking: 42%Fable 5.1 on τ³-Banking: 47%47%CritPtFable 5 on CritPt: 29%Opus 5 on CritPt: 29%Fable 5.1 on CritPt: 30%30%
The seven evaluations in Intelligence Index v4.1.1 that are scored as percentage accuracy, so they share one axis honestly. Labelled values are Fable 5.1. Three more evaluations use their own scales and appear in the table below instead: GDPval-AA v2 is a rescaled Elo, AA-Omniscience Index runs from -100 to 100, and AA-Briefcase is an Elo rating. Figures are read off the Artificial Analysis comparison charts and rounded to the nearest point, so treat one-point differences as noise - where AA publishes a precise number it matches, putting Fable 5.1 at 91.4% on Terminal-Bench v2.1, the highest score on that evaluation.
EvaluationScaleFable 5.1Fable 5Opus 5
GPQA Diamond scientific reasoning%949393
Terminal-Bench v2.1 agentic coding%918589
AA-LCR long-context reasoning%807776
SciCode coding%626056
Humanity's Last Exam reasoning & knowledge%595555
𝜏³-Banking agentic tool use%473842
CritPt physics reasoning%302929
GDPval-AA v2 real-world work tasksindex686166
AA-Omniscience knowledge reliability-100 to 100434237
AA-Briefcase agentic knowledge workElo169415171685

Is Fable 5.1 actually cheaper?

Anthropic says 5.1 costs about 25% less than Fable 5 for typical workloads, and up to 45% less for agentic work. Artificial Analysis lists it as more expensive than Fable 5. We checked, and both are right - they are measuring different things.

060120180Fable 5 = 100Blended price per 1M tokensBlended price per 1M tokens: 93 indexed ($7.70 → $7.17)93$7.70 → $7.17Cost per Intelligence Index taskCost per Intelligence Index task: 118 indexed ($3.14 → $3.69)118$3.14 → $3.69Output tokens to run the indexOutput tokens to run the index: 169 indexed (83M → 140M)16983M → 140M
Everything indexed against Fable 5 = 100. The unit price of tokens did fall. The bill for a unit of work went up, because 5.1 emits far more tokens to do it.
The price cut is real, and Artificial Analysis's own numbers confirm it. Anthropic cut cache reads by 75%, from $1.00 to $0.25 per million tokens, leaving input and output at $10 and $50. AA's blended price for Fable 5.1 at Anthropic is $7.17 per 1M tokens against $7.70 for Fable 5 - and at AA's 7:2:1 cache-to-input-to-output ratio, $7.17 back-solves to a cache read of exactly $0.25. So the two sources agree on the price; the cut is not in dispute.
What AA is actually flagging is verbosity, not price. Fable 5.1 generated 140M output tokens to complete the Intelligence Index, against 83M for Fable 5 - 69% more. Output tokens are the expensive kind and the index barely uses caching, so the total cost of running it rose from $5,455 to $8,523 and cost per task rose from $3.14 to $3.69. On that measure 5.1 is 18% more expensive per unit of work, and AA's "particularly expensive" verdict follows from it.

Which figure applies to you comes down to two things. If your workload is cache-heavy - long-running agents re-reading a large context, which is what Claude Code does all day - cache reads dominate your bill and the 75% cut lands close to Anthropic's numbers. If your workload is generation-heavy and cache-light, closer to how AA runs its index, you may well pay more per finished task than you did on Fable 5, because the extra reasoning tokens outweigh the cache saving. Note also that AA's 7:2:1 blend puts output at 65% of the blended figure, which is why a 75% cut to cache reads only moves that number 7%: the blend understates the saving for agentic work just as the index overstates the cost.

DeepSWE has not scored it yet

No Fable 5.1 entry as of September 2, 2026. The DeepSWE leaderboard was last updated August 26, 2026 - before Fable 5.1 shipped - so it carries no 5.1 result. For reference, its current top three over 113 tasks are claude-opus-5 [max] at 74% ±4% ($11.84 per task), gpt-5.6-sol [max] at 73% ±3% ($6.46), and claude-fable-5 [max] at 70% ±4% ($21.63). Fable 5 is by far the most expensive model per task in that top group, which is the gap the cache-read cut is aimed at. We will add 5.1 here once DeepSWE publishes it.

What the numbers do not say

Artificial Analysis measures Fable 5.1 with production safeguards on and fallbacks enabled, so a blocked request is answered by a different Claude model rather than scored zero - the index reflects the deployed system, not the raw model. And no third party has yet measured the thing Anthropic leads with: Terminal-Bench-Science 0.1, where Anthropic reports 52.6% for 5.1 against 24.7% for Fable 5. That is the largest claimed jump in the release and it is currently unverified outside Anthropic's own harness.