← Overall rankings

LMArena Hard prompts arena · published July 20, 2026

Anthropic leads the LMArena hard prompts arena.

As of July 20, 2026, Anthropic leads the LMArena hard prompts arena with claude-fable-5 at 1533.4 across 6,523 battles, according to LMArena.

Current through July 20, 2026. 167 published snapshots. 382 of 386 models resolved (4 held for exact-match review).

LMArena hard prompts arena rating, published July 20, 2026. Top 25 of 324 rows.
RankLMArena ratingOrganizationModelReported intervalGap to leaderBattles
011533.4AnthropicClaude Fable 51525.3 to 1541.6Leader6,523
021532.2Anthropicclaude-opus-4-6-thinking1527.6 to 1536.8-1.238,977
031526.7Anthropicclaude-opus-4-61522.2 to 1531.2-6.742,370
041525.9Anthropicclaude-opus-4-7-thinking1520.9 to 1531.0-7.533,058
051519.1AnthropicClaude Opus 4.71514.1 to 1524.2-14.333,785
061514.3Anthropicclaude-opus-4-8-thinking1508.1 to 1520.4-19.119,745
071513.6Metamuse-spark-1.11504.5 to 1522.7-19.84,692
081506.4Googlegemini-3.1-pro-preview1502.0 to 1510.7-27.053,341
091506.1OpenAIgpt-5.6-sol-xhigh1495.8 to 1516.4-27.33,559
101505.8Anthropicclaude-opus-4-81499.7 to 1511.9-27.620,263
111504.1MetaMuse Spark1497.1 to 1511.2-29.38,707
121504.0Anthropicclaude-sonnet-4-61499.3 to 1508.8-29.436,123
131504.0GoogleGemini 3 Pro1499.0 to 1509.0-29.422,448
141500.4OpenAIgpt-5.5-high1495.1 to 1505.6-33.029,190
151499.3Anthropicclaude-opus-4-5-20251101-thinking-32k1494.2 to 1504.5-34.119,755
161498.7OpenAIgpt-5.4-high1494.0 to 1503.5-34.737,471
171497.8Anthropicclaude-opus-4-5-202511011493.7 to 1501.9-35.641,450
181496.8OpenAIgpt-5.2-chat-latest-202602101491.7 to 1501.9-36.621,642
191496.3Alibabaqwen3.7-max-preview1484.1 to 1508.6-37.12,526
201494.8Googlegemini-3.5-flash-medium1487.5 to 1502.1-38.68,796
211494.5Xiaomimimo-v2.5-pro1489.2 to 1499.8-38.926,993
221493.5OpenAIgpt-5.51488.3 to 1498.7-39.930,515
231492.4Googlegemini-3-flash1486.7 to 1498.0-41.016,565
241491.9Z.aiglm-5.11486.5 to 1497.4-41.519,738
251491.5Googlegemini-3.5-flash-high1483.6 to 1499.4-41.96,779

Frontier race

Category Kings: Hard prompts arena frontier by organization

Anthropic holds the hard prompts arena frontier at 1533 as of July 20, 2026.

Hard prompts arena frontier ratings by organization over timeAnthropic holds the hard prompts arena frontier at 1533 as of July 20, 2026. 10 organizations are tracked from August 28, 2024 to July 20, 2026.114012501360147015802024-082025-022025-072025-122026-032026-07Anthropic 1533Meta 1514Google 1506OpenAI 1506Alibaba 1496Xiaomi 1495Z.ai 1492xAI 1490DeepSeek 1480Mistral AI 1445

Method: Frontier is the maximum resolved LMArena rating per organization on each publish date. Organizations limited to the union of the seven tracked homepage providers and organizations in the latest board top 25 by officialRank. Quarantined dates are excluded, and this category series keeps its own method epoch and is never joined to the legacy archive.

Hard prompts arena frontier by organization: latest resolved score per organization, from the reviewed LMArena category artifact.
OrganizationLatest modelLatest LMArena ratingFirst tracked date
Anthropicanthropic/claude-fable-51533.4August 28, 2024
Metameta/muse-spark-1.11513.6August 28, 2024
Googlegoogle/gemini-3.1-pro-preview1506.4August 28, 2024
OpenAIopenai/gpt-5.6-sol-xhigh1506.1August 28, 2024
Alibabaalibaba/qwen3.7-max-preview1496.3August 28, 2024
Xiaomixiaomi/mimo-v2.5-pro1494.5December 23, 2025
Z.aizai/glm-5.11491.9August 28, 2024
xAIxai/grok-4.20-beta-0309-reasoning1489.8August 28, 2024
DeepSeekdeepseek/deepseek-v4-pro1480.2August 28, 2024
Mistral AImistral/mistral-medium-3.51445.4August 28, 2024
Anthropic leads the hard prompts arena frontier at 1533.4 on July 20, 2026. Frontier is the highest resolved score per organization per publish date. Quarantined dates are excluded and this category is never joined to the legacy archive.

Definitions

What these terms mean.

Hard prompts arena
The LMArena style-controlled text battles filtered to prompts the source classifies as hard. Difficulty classification belongs to LMArena, not Model Gauntlet.
Frontier by organization
The highest resolved score for each organization on each publish date. It tracks how far each lab has pushed this category, not an average of all its models.
Confidence interval
The reported 95% range LMArena publishes with each rating. Overlapping intervals mean this snapshot does not separate those ranks decisively.

Hard prompts arena questions

Which model leads the Hard prompts arena right now?

According to the LMArena hard prompts board published July 20, 2026, Anthropic leads with claude-fable-5 at 1533.4, across 6,523 battles. That is 1.2 points ahead of the next board entry inside this one snapshot.

Is this a Model Gauntlet hard prompts ranking?

No. This board mirrors one named source, the LMArena hard prompts arena, published July 20, 2026. Model Gauntlet publishes no composite or consensus score and runs no evaluation of its own.

How many models and organizations are in the hard prompts board?

The reviewed artifact covers 386 models across 61 organizations, with 382 models resolved to exact identities and 4 held for review. The history spans 167 snapshots from August 28, 2024 to July 20, 2026.

Why do some scores have overlapping confidence intervals?

LMArena publishes a 95% confidence interval with each rating. Where two intervals overlap, this single snapshot does not separate those ranks decisively, so treat small hard prompts gaps with caution.