Capability progress is outrunning task-cost improvement
Higher reasoning effort is lifting capability scores and task costs together. Artificial Analysis measures Fable 5.1 at 66 with a 20% higher task bill, and Gemini 3.8 Flash at 59 and $0.58 per task, about 40% above its predecessor despite unchanged token prices. Epoch’s separate frontier fit of 14 ECI points a year versus 6 measures capability progress, not equivalent cost savings.
updates 2026-09-01 · prior evidence ↗
Claude Fable 5.1 tops the Artificial Analysis Intelligence Index but costs 20% more per task than Fable 5 despite a 75% cache read price cut We supported @AnthropicAI with pre-release evaluation of Claude Fable 5.1. At max effort it scores 66 on the Artificial Analysis…
Claude Fable 5.1 tops the Artificial Analysis Intelligence Index but costs 20% more per task than Fable 5 despite a 75% cache read price cut We supported @AnthropicAI with pre-release evaluation of Claude Fable 5.1. At max effort it scores 66 on the Artificial Analysis Intelligence Index, the highest score we have measured, ahead of Claude Opus 5 (max, 63), Claude Fable 5 (max, 62), GPT-5.6 Sol (max, 61) and Grok 4.6 (high, 61). We evaluated the model with Anthropic's ‘default’ server-side fallback, which routes safety-flagged requests to Claude Opus 4.8 or Claude Opus 5; fallback served ~4% of output tokens across the Intelligence Index. Key takeaways ➤ Frontier Intelligence with improvements across benchmarks: Fable 5.1 gains +4 points on the Intelligence Index over Fable 5. On HLE, Fable 5.1 scores 59.1%, ahead of the previous best of 55.5% from Claude Fable 5. It posts the narrowly highest scores we’ve seen on Terminal-Bench v2.1 (91.4%) and SciCode (62.0%), and on τ³-Banking it gains 9 points over Fable 5 ➤ 75% cache read price cut, but Fable 5.1 still costs more per task: Anthropic has cut the cache read price from $1 to $0.25 per 1M cached input tokens, with standard pricing unchanged at $10/$50 per 1M input/output tokens. Fable 5.1 (max) costs $3.76 per Intelligence Index task, 20% more than Fable 5 (max), because it uses ~1.7x the output tokens. The cache cut saves ~$1.40 per task, concentrated in the agentic evaluations where the majority of input tokens are cache reads. At xhigh effort Fable 5.1 scores 65 at $2.72 per task, $1.04 less than max, but still above Claude Opus 5 (max, 63) at $2.34 ➤ Claude Fable 5.1 holds the upper end of the Intelligence vs Output Tokens per Task Pareto frontier: every model variant scoring higher than GPT-5.6 Sol (medium) on the Intelligence Index is matched or beaten by a Fable 5.1 effort level on both intelligence and token usage ➤ Highest scores on agentic work tasks, but effectively tied with Opus 5: Fable 5.1 sets the highest scores we have measured on GDPval-AA v2 (1,853 Elo, +130 over Fable 5) and AA-Briefcase (1,694 Elo, +122 over Fable 5), our agentic knowledge work evaluations. Against Claude Opus 5 the GDPval-AA v2 lead is within the confidence interval and AA-Briefcase (1,685) is effectively tied, with Fable 5.1 ahead on analytical quality and rubric correctness, but behind on presentation Other model details: ➤ Context window: 1 million tokens, supporting image and text inputs as with Anthropic’s other recent launches ➤ Pricing: Fable 5.1 retains the $10/$50/$12.5 input, output, and cache write prices per million tokens from Fable 5, but cache hits have been reduced to $0.25 per million tokens, a 75% relative reduction from before that will materially reduce agentic workload costs
discussion1 selected reply
Google has released Gemini 3.8 Flash, its fourth Flash model in under four months - it scores 59 on the Artificial Analysis Intelligence Index and reaches the Intelligence vs. Cost per Task Pareto frontier @GoogleDeepMind released Gemini 3.8 Flash today. With high reasoning,…
Google has released Gemini 3.8 Flash, its fourth Flash model in under four months - it scores 59 on the Artificial Analysis Intelligence Index and reaches the Intelligence vs. Cost per Task Pareto frontier @GoogleDeepMind released Gemini 3.8 Flash today. With high reasoning, it scores 59 on the Artificial Analysis Intelligence Index, up 3 points from Gemini 3.7 Flash and on par with sub-maximum reasoning efforts of GPT-5.6 Sol (xhigh, 59) and Grok 4.6 (medium, 59) Matching Gemini 3.7 Flash’s discounted pricing until the end of the year ($0.75/$3.75 per million input/output tokens), Gemini 3.8 Flash sits on the Intelligence vs. Cost per Task Pareto frontier at $0.58 per task. This is comparable to GPT-5.6 Terra (max, $0.53), but ~40% higher than its predecessor, driven by a 30% increase in average output tokens per task to 48k and increased turns on agentic evaluations Key benchmarking results across Gemini 3.8 Flash’s three reasoning levels: ➤ 3 point Intelligence Index improvement: Gemini 3.8 Flash (high) scores 59 on the Artificial Analysis Intelligence Index, up 3 points from Gemini 3.7 Flash (high, 56). With medium reasoning it scores 57, matching GPT-5.6 Terra (max, 57) and Muse Spark 1.2 (xhigh, 57). With low reasoning it scores 52, matching Gemini 3.6 Flash (high, 52), at 30% lower Cost per Task and roughly a third of the Time per Task ➤ Agentic capability improvements: Gemini 3.8 Flash’s 3 point improvement on the Artificial Analysis Intelligence Index is primarily driven by stronger performance on agentic evaluations such as 𝜏³-Banking (tool use), Terminal-Bench v2.1 (coding) and GDPval-AA v2 (real-world tasks). The largest improvement is on 𝜏³-Banking, where it gains 12 points over Gemini 3.7 Flash to score 45% ➤ Pareto frontier on Intelligence vs. Cost per Task: Gemini 3.8 Flash (high) costs $0.58 per Intelligence Index task, making it the cheapest model at its level of intelligence. This is up ~40% from Gemini 3.7 Flash ($0.40) despite unchanged per-token pricing, driven by a 30% increase in output tokens per task and more turns on agentic evaluations. Cost per Task falls to $0.41 with medium reasoning and $0.24 with low reasoning ➤ Output speeds remain fast, but Time per Task increases: On high reasoning, Gemini 3.8 Flash averages ~300 output tokens per second and a Time per Task of 2.5 minutes, slightly faster than GPT-5.6 Luna (max, 2.6 minutes) and GPT-5.6 Terra (max, 3.3). Compared to Gemini 3.7 Flash, higher token usage increases Time per Task from 2.2 minutes to 2.5 minutes, and puts it behind Claude Fable 5.1 (medium, 2.1 minutes). On low reasoning, Time per Task falls to 0.8 minutes, placing Gemini 3.8 Flash on the Intelligence vs. Time per Task Pareto frontier Key model details: ➤ Context Window: 1M tokens, unchanged from Gemini 3.7 Flash ➤ Multimodality: Text, image, video, and speech input, with text output ➤ Pricing: $0.75/$3.75 per 1M input/output tokens through the end of the year, matching Gemini 3.7 Flash’s current discounted pricing. $1.50/$7.50 per 1M input/output tokens at standard pricing. Cached input tokens retain the same 90% discount

Claude Fable 5.1's five effort settings span 11x in output token usage, from 13.1M at low effort to 143.7M at max, and score from 58 to 66 on the Artificial Analysis Intelligence Index. Across effort levels Fable 5.1 sits on the Intelligence vs Output Tokens frontier, but its floor is higher than the GPT-5.6 family's: GPT-5.6 Sol (medium) uses marginally fewer tokens than Fable 5.1 (low), 12M against 13.1M.