Claude Fable 5
- LMArena rank
- #1 · 1506.8
- Agent Arena
- #1 · Claude Fable 5 (High)
- OpenRouter usage
- #24 · 449.9B tokens/week
- OpenRouter price
- $10.00 in / $50.00 out per 1M
MODEL GAUNTLET RESEARCH DESK
SOURCE-ATTRIBUTED AI RANKING AGGREGATOR EDITORIALLY INDEPENDENT
Anthropic leads the July 16, 2026 LMArena text ranking.
The current frontier
An editorial watchlist of the models in active frontier discussion, audited July 22, 2026: every pick appears on at least one current reviewed board. Selection is ours; every number on a card is one named source's own value with its own scale, every card carries the same stat rows, and a board that does not list a model says so. Nothing is blended.
LMArena rating · style-controlled text · published July 20, 2026
Bars start at zero on LMArena's own scale, so close ratings look close. Hover or focus a bar for the reported interval and battle count. The cards below carry the same values as text.
7 Epoch hub benchmarks carry this model. See the benchmark section.
Sources: LMArena category artifact published July 20, 2026, Aider polyglot updated May 22, 2026, EQ-Bench 3 updated May 10, 2026, and the Epoch AI Benchmarking Hub retrieved July 19, 2026. Each carries its receipt in its labeled section below. Rows named with a scaffold are system results, not bare-model results.
Chatbot Arena archive
Meta highlighted. 827 to 1288 published rating; sampled peak 1396.
Line ends, top to bottom:Google 1466Mistral 1433xAI 1431OpenAI 1429DeepSeek 1424Anthropic 1417Meta 1288
Anthropic's Claude Fable 5 holds the highest provider-frontier point estimate in LMArena's July 16, 2026 text style-control overall snapshot, with Meta's Muse Spark 1.1 in second place.
· Leaderboard update
· Model release
· Model release
· Model release
· Model release
Current LMArena snapshot
These values come from LMArena’s text style-control overall table published July 16, 2026. The legacy archive remains separate.
Claude Fable 5 is the highest-scoring model in LMArena’s July 16, 2026 snapshot, with a reported interval of 1500.2 to 1514.7.
Meta’s Muse Spark 1.1 ranks second at 1493.1, 14.4 points behind Anthropic.
LMArena reports 8,817 battles for Claude Fable 5 in this published snapshot.
Each model is matched to the organization listed in LMArena’s published data, avoiding guesses based on model names.
Point-in-time standings
We take each organization’s highest LMArena rating in this published snapshot. Confidence intervals show where close results may not represent a meaningful difference.
| Source rank | Model | Organization | LMArena rating | Reported interval | Battles |
|---|---|---|---|---|---|
| 01 | Claude Fable 5 | 1507.5 | 1500.2–1514.7 | 8,817 | |
| 02 | Muse Spark 1.1 | 1493.1 | 1484.9–1501.3 | 5,732 | |
| 03 | GPT-5.6 Sol X-High | 1486.3 | 1476.8–1495.8 | 4,113 | |
| 04 | Gemini 3 Pro | 1485.8 | 1481.9–1489.6 | 41,283 | |
| 05 | Grok 4.20 Beta 1 | 1474.2 | 1469.5–1478.8 | 26,844 | |
| 06 | DeepSeek V4 Pro | 1457.5 | 1453.1–1462.0 | 43,062 | |
| 07 | Mistral Medium 3.5 | 1427.3 | 1420.7–1433.9 | 11,017 |
LMArena category boards
Each board is LMArena's own category rating. Open one board at a time; every board links to its full page with intervals and history.
Aider polyglot · code editing
Aider ranks full model-and-editor system configurations on code editing, not bare models. Pass rate 2 is Aider's own metric, and these rows are never compared against LMArena ratings.
Agent Arena · agentic tasks
LMArena's Agent Arena ranks models on real agentic sessions and published this board July 20, 2026.IPS is this arena's own score with its own scale. Rows labeled with a reasoning configuration are system results for that configuration, and a score at or below zero draws no bar.
EQ-Bench 3 · emotional intelligence
EQ-Bench 3 publishes its own Elo from judged role-play scenarios. The scale belongs to this source alone and shares nothing with LMArena numbers.
Epoch AI Benchmarking Hub
Every card is one benchmark's own top scores from Epoch AI's published hub data. Rows labeled with a scaffold, such as a coding agent or edit format, are system results, not bare-model results. Fraction-unit scores are shown as percentages.
Aider Polyglot leaderboard · 69 resolved scores
Epoch AI Benchmarking Hub · 54 resolved scores
Epoch AI Benchmarking Hub · 98 resolved scores
Epoch AI Benchmarking Hub · 69 resolved scores
Epoch AI Benchmarking Hub · 176 resolved scores
Epoch AI Benchmarking Hub · 106 resolved scores
Epoch AI Benchmarking Hub · 150 resolved scores
Epoch AI Benchmarking Hub · 62 resolved scores
Epoch AI Benchmarking Hub · 33 resolved scores
Terminal-Bench leaderboard · 178 resolved scores
OpenRouter platform view
OpenRouter's own usage rankings, model listings, and routed prices. Usage is the platform's weekly routed token volume,retrieved July 22, 2026: a live popularity signal that shows which models developers route work through right now, not a quality score. The platform's embedded third-party benchmark data is excluded at collection with a receipt and never appears on this site.
OpenRouter's routed token volume, week view. Volume is not a quality signal.
By OpenRouter's listing date. A listing date is not a release date.
Input price per million tokens, cheapest first. Price says nothing about quality.
From the archive
827 → 1396 → 1288
This history includes Meta’s own models and excludes third-party models built from Llama.
Chart lab
Methodology
Explore Model Gauntlet
Rankings, model records, research reports, and interactive history. Every claim is linked to its source, with clear evidence boundaries.
Research and commercial inquiries
Research questions, source submissions, corrections, or partnerships. Commercial relationships never influence a score, ranking, or editorial conclusion.
Direct answers
Model Gauntlet publishes sourced AI model standings and a separately labeled historical race archive. Current and legacy metrics are never blended.
The current provider table uses LMArena's published text style-control overall snapshot, dated July 16, 2026. It includes confidence intervals and battle counts.
No. The current LMArena rating and the legacy May 2023 to August 2025 archive use incompatible methodology and inclusion rules. No cross-window change is calculated.
Race alerts
Join the waitlist for sourced updates on leaderboard changes, major model releases, and new research. The briefing launches after a final editorial review; early subscribers get the first edition.