Everything on this page comes from Artificial Analysis and the DeepSWE leaderboard - independent third parties, not Anthropic's own eval numbers. Scores move; the capture date above is when these were read.
The headline
Intelligence Index
66#1 of 192 models in its class, on Intelligence Index v4.1.1.
Cost per index task
$3.69#84 of 192 - and up from $3.14 on Fable 5. See the pricing section below.
Output speed
66.4 tok/s#85 of 192 - slower than the 71 tok/s median.
Verbosity
140M tokensOutput tokens across the whole index, against a 71M median. The most verbose model measured.
Artificial Analysis summarises it as "amongst the leading models in intelligence, but particularly expensive when comparing to other models of similar price... also slower than average and very verbose." Running the full index on Fable 5.1 cost them $8,523.16.
Against the rest of the frontier
Artificial Analysis Intelligence Index v4.1.1, top models by score. Fable 5.1 takes the top two slots, but the four-point lead over Fable 5 is worth reading next to the cost column in the table below: GPT-5.6 Sol is within five points at roughly a quarter of the cost per task. Source: Artificial Analysis leaderboard.
Model
Intelligence Index
Cost per task
Median tok/s
Context
Claude Fable 5.1 (max)
66
$3.69
66
1M
Claude Fable 5.1 (xhigh)
65
$2.65
59
1M
Claude Opus 5 (max)
63
$2.34
52
1M
Claude Fable 5 (max)
62
$3.14
65
1M
GPT-5.6 Sol (max)
61
$0.95
74
1M
Grok 4.6 (high)
61
$0.94
53
500k
Kimi K3 (max)
60
$0.84
38
1.05M
GLM-5.3 (max)
60
$0.68
70
1M
Intelligence by effort level
Fable 5.1 is measured at every effort setting, and the two panels below do not rise at the same rate. Going from low to max buys eight index points and costs 4.8× as much per task.
Left: Intelligence Index by effort level. Right: cost per index task, same effort levels, plotted on its own axis because the two measures share no scale. Low effort already scores 58 - within eight points of max - for 21% of the cost, which is the practical argument against defaulting everything to max. All levels are the "with fallback" configurations, meaning safeguard-blocked requests are answered by another Claude model rather than scored zero.
Per-evaluation breakdown
Fable 5.1 leads or ties Fable 5 and Opus 5 on every evaluation in the index. The widest gaps are agentic: 𝜏³-Banking (tool use) and Terminal-Bench v2.1 (terminal coding).
Fable 5.1 Fable 5 Opus 5
The seven evaluations in Intelligence Index v4.1.1 that are scored as percentage accuracy, so they share one axis honestly. Labelled values are Fable 5.1. Three more evaluations use their own scales and appear in the table below instead: GDPval-AA v2 is a rescaled Elo, AA-Omniscience Index runs from -100 to 100, and AA-Briefcase is an Elo rating. Figures are read off the Artificial Analysis comparison charts and rounded to the nearest point, so treat one-point differences as noise - where AA publishes a precise number it matches, putting Fable 5.1 at 91.4% on Terminal-Bench v2.1, the highest score on that evaluation.
Evaluation
Scale
Fable 5.1
Fable 5
Opus 5
GPQA Diamond scientific reasoning
%
94
93
93
Terminal-Bench v2.1 agentic coding
%
91
85
89
AA-LCR long-context reasoning
%
80
77
76
SciCode coding
%
62
60
56
Humanity's Last Exam reasoning & knowledge
%
59
55
55
𝜏³-Banking agentic tool use
%
47
38
42
CritPt physics reasoning
%
30
29
29
GDPval-AA v2 real-world work tasks
index
68
61
66
AA-Omniscience knowledge reliability
-100 to 100
43
42
37
AA-Briefcase agentic knowledge work
Elo
1694
1517
1685
Is Fable 5.1 actually cheaper?
Anthropic says 5.1 costs about 25% less than Fable 5 for typical workloads, and up to 45% less for agentic work. Artificial Analysis lists it as more expensive than Fable 5. We checked, and both are right - they are measuring different things.
Everything indexed against Fable 5 = 100. The unit price of tokens did fall. The bill for a unit of work went up, because 5.1 emits far more tokens to do it.
The price cut is real, and Artificial Analysis's own numbers confirm it.
Anthropic cut cache reads by 75%, from $1.00 to $0.25 per million tokens, leaving input and output at $10 and $50. AA's blended price for Fable 5.1 at Anthropic is $7.17 per 1M tokens against $7.70 for Fable 5 - and at AA's 7:2:1 cache-to-input-to-output ratio, $7.17 back-solves to a cache read of exactly $0.25. So the two sources agree on the price; the cut is not in dispute.
What AA is actually flagging is verbosity, not price.
Fable 5.1 generated 140M output tokens to complete the Intelligence Index, against 83M for Fable 5 - 69% more. Output tokens are the expensive kind and the index barely uses caching, so the total cost of running it rose from $5,455 to $8,523 and cost per task rose from $3.14 to $3.69. On that measure 5.1 is 18% more expensive per unit of work, and AA's "particularly expensive" verdict follows from it.
Which figure applies to you comes down to two things. If your workload is cache-heavy - long-running agents re-reading a large context, which is what Claude Code does all day - cache reads dominate your bill and the 75% cut lands close to Anthropic's numbers. If your workload is generation-heavy and cache-light, closer to how AA runs its index, you may well pay more per finished task than you did on Fable 5, because the extra reasoning tokens outweigh the cache saving. Note also that AA's 7:2:1 blend puts output at 65% of the blended figure, which is why a 75% cut to cache reads only moves that number 7%: the blend understates the saving for agentic work just as the index overstates the cost.
DeepSWE has not scored it yet
No Fable 5.1 entry as of September 2, 2026.
The DeepSWE leaderboard was last updated August 26, 2026 - before Fable 5.1 shipped - so it carries no 5.1 result. For reference, its current top three over 113 tasks are claude-opus-5 [max] at 74% ±4% ($11.84 per task), gpt-5.6-sol [max] at 73% ±3% ($6.46), and claude-fable-5 [max] at 70% ±4% ($21.63). Fable 5 is by far the most expensive model per task in that top group, which is the gap the cache-read cut is aimed at. We will add 5.1 here once DeepSWE publishes it.
What the numbers do not say
Artificial Analysis measures Fable 5.1 with production safeguards on and fallbacks enabled, so a blocked request is answered by a different Claude model rather than scored zero - the index reflects the deployed system, not the raw model. And no third party has yet measured the thing Anthropic leads with: Terminal-Bench-Science 0.1, where Anthropic reports 52.6% for 5.1 against 24.7% for Fable 5. That is the largest claimed jump in the release and it is currently unverified outside Anthropic's own harness.