LMArena Hard prompts arena · published July 20, 2026
Anthropic leads the LMArena hard prompts arena.
As of July 20, 2026, Anthropic leads the LMArena hard prompts arena with claude-fable-5 at 1533.4 across 6,523 battles, according to LMArena.
Current through July 20, 2026. 167 published snapshots. 382 of 386 models resolved (4 held for exact-match review).
| Rank | LMArena rating | Organization | Model | Reported interval | Gap to leader | Battles |
|---|---|---|---|---|---|---|
| 01 | 1533.4 | Anthropic | Claude Fable 5 | 1525.3 to 1541.6 | Leader | 6,523 |
| 02 | 1532.2 | Anthropic | claude-opus-4-6-thinking | 1527.6 to 1536.8 | -1.2 | 38,977 |
| 03 | 1526.7 | Anthropic | claude-opus-4-6 | 1522.2 to 1531.2 | -6.7 | 42,370 |
| 04 | 1525.9 | Anthropic | claude-opus-4-7-thinking | 1520.9 to 1531.0 | -7.5 | 33,058 |
| 05 | 1519.1 | Anthropic | Claude Opus 4.7 | 1514.1 to 1524.2 | -14.3 | 33,785 |
| 06 | 1514.3 | Anthropic | claude-opus-4-8-thinking | 1508.1 to 1520.4 | -19.1 | 19,745 |
| 07 | 1513.6 | Meta | muse-spark-1.1 | 1504.5 to 1522.7 | -19.8 | 4,692 |
| 08 | 1506.4 | gemini-3.1-pro-preview | 1502.0 to 1510.7 | -27.0 | 53,341 | |
| 09 | 1506.1 | OpenAI | gpt-5.6-sol-xhigh | 1495.8 to 1516.4 | -27.3 | 3,559 |
| 10 | 1505.8 | Anthropic | claude-opus-4-8 | 1499.7 to 1511.9 | -27.6 | 20,263 |
| 11 | 1504.1 | Meta | Muse Spark | 1497.1 to 1511.2 | -29.3 | 8,707 |
| 12 | 1504.0 | Anthropic | claude-sonnet-4-6 | 1499.3 to 1508.8 | -29.4 | 36,123 |
| 13 | 1504.0 | Gemini 3 Pro | 1499.0 to 1509.0 | -29.4 | 22,448 | |
| 14 | 1500.4 | OpenAI | gpt-5.5-high | 1495.1 to 1505.6 | -33.0 | 29,190 |
| 15 | 1499.3 | Anthropic | claude-opus-4-5-20251101-thinking-32k | 1494.2 to 1504.5 | -34.1 | 19,755 |
| 16 | 1498.7 | OpenAI | gpt-5.4-high | 1494.0 to 1503.5 | -34.7 | 37,471 |
| 17 | 1497.8 | Anthropic | claude-opus-4-5-20251101 | 1493.7 to 1501.9 | -35.6 | 41,450 |
| 18 | 1496.8 | OpenAI | gpt-5.2-chat-latest-20260210 | 1491.7 to 1501.9 | -36.6 | 21,642 |
| 19 | 1496.3 | Alibaba | qwen3.7-max-preview | 1484.1 to 1508.6 | -37.1 | 2,526 |
| 20 | 1494.8 | gemini-3.5-flash-medium | 1487.5 to 1502.1 | -38.6 | 8,796 | |
| 21 | 1494.5 | Xiaomi | mimo-v2.5-pro | 1489.2 to 1499.8 | -38.9 | 26,993 |
| 22 | 1493.5 | OpenAI | gpt-5.5 | 1488.3 to 1498.7 | -39.9 | 30,515 |
| 23 | 1492.4 | gemini-3-flash | 1486.7 to 1498.0 | -41.0 | 16,565 | |
| 24 | 1491.9 | Z.ai | glm-5.1 | 1486.5 to 1497.4 | -41.5 | 19,738 |
| 25 | 1491.5 | gemini-3.5-flash-high | 1483.6 to 1499.4 | -41.9 | 6,779 |
Frontier race
Category Kings: Hard prompts arena frontier by organization
Anthropic holds the hard prompts arena frontier at 1533 as of July 20, 2026.
Method: Frontier is the maximum resolved LMArena rating per organization on each publish date. Organizations limited to the union of the seven tracked homepage providers and organizations in the latest board top 25 by officialRank. Quarantined dates are excluded, and this category series keeps its own method epoch and is never joined to the legacy archive.
| Organization | Latest model | Latest LMArena rating | First tracked date |
|---|---|---|---|
| Anthropic | anthropic/claude-fable-5 | 1533.4 | August 28, 2024 |
| Meta | meta/muse-spark-1.1 | 1513.6 | August 28, 2024 |
| google/gemini-3.1-pro-preview | 1506.4 | August 28, 2024 | |
| OpenAI | openai/gpt-5.6-sol-xhigh | 1506.1 | August 28, 2024 |
| Alibaba | alibaba/qwen3.7-max-preview | 1496.3 | August 28, 2024 |
| Xiaomi | xiaomi/mimo-v2.5-pro | 1494.5 | December 23, 2025 |
| Z.ai | zai/glm-5.1 | 1491.9 | August 28, 2024 |
| xAI | xai/grok-4.20-beta-0309-reasoning | 1489.8 | August 28, 2024 |
| DeepSeek | deepseek/deepseek-v4-pro | 1480.2 | August 28, 2024 |
| Mistral AI | mistral/mistral-medium-3.5 | 1445.4 | August 28, 2024 |
Definitions
What these terms mean.
- Hard prompts arena
- The LMArena style-controlled text battles filtered to prompts the source classifies as hard. Difficulty classification belongs to LMArena, not Model Gauntlet.
- Frontier by organization
- The highest resolved score for each organization on each publish date. It tracks how far each lab has pushed this category, not an average of all its models.
- Confidence interval
- The reported 95% range LMArena publishes with each rating. Overlapping intervals mean this snapshot does not separate those ranks decisively.
Follow the evidence
Compare this board against the overall LMArena rankings, shortlist candidates with the Reasoning and research guide,read the Anthropic lab file, and see how rankings are selected. Model profiles on this board: Claude Fable 5, Claude Opus 4.7, Muse Spark.
Hard prompts arena questions
Which model leads the Hard prompts arena right now?
According to the LMArena hard prompts board published July 20, 2026, Anthropic leads with claude-fable-5 at 1533.4, across 6,523 battles. That is 1.2 points ahead of the next board entry inside this one snapshot.
Is this a Model Gauntlet hard prompts ranking?
No. This board mirrors one named source, the LMArena hard prompts arena, published July 20, 2026. Model Gauntlet publishes no composite or consensus score and runs no evaluation of its own.
How many models and organizations are in the hard prompts board?
The reviewed artifact covers 386 models across 61 organizations, with 382 models resolved to exact identities and 4 held for review. The history spans 167 snapshots from August 28, 2024 to July 20, 2026.
Why do some scores have overlapping confidence intervals?
LMArena publishes a 95% confidence interval with each rating. Where two intervals overlap, this single snapshot does not separate those ranks decisively, so treat small hard prompts gaps with caution.