← Overall rankings

LMArena Instruction following arena · published July 20, 2026

Anthropic leads the LMArena instruction following arena.

As of July 20, 2026, Anthropic leads the LMArena instruction following arena with claude-fable-5 at 1512.9 across 3,433 battles, according to LMArena.

Current through July 20, 2026. 145 published snapshots. 378 of 382 models resolved (4 held for exact-match review).

LMArena instruction following arena rating, published July 20, 2026. Top 25 of 324 rows.
RankLMArena ratingOrganizationModelReported intervalGap to leaderBattles
011512.9AnthropicClaude Fable 51502.3 to 1523.6Leader3,433
021512.8Anthropicclaude-opus-4-6-thinking1507.1 to 1518.6-0.119,376
031503.8Anthropicclaude-opus-4-7-thinking1497.6 to 1510.0-9.116,905
041498.9Anthropicclaude-opus-4-61493.4 to 1504.4-14.021,316
051494.2Anthropicclaude-opus-4-8-thinking1487.0 to 1501.5-18.710,141
061492.9AnthropicClaude Opus 4.71486.9 to 1499.0-20.017,352
071483.9Anthropicclaude-opus-4-5-20251101-thinking-32k1477.1 to 1490.6-29.09,450
081483.5OpenAIgpt-5.6-sol-xhigh1469.1 to 1497.9-29.41,787
091481.2Googlegemini-3.1-pro-preview1476.0 to 1486.4-31.727,250
101478.1Metamuse-spark-1.11465.6 to 1490.5-34.82,443
111477.9Anthropicclaude-opus-4-81470.6 to 1485.2-35.010,394
121477.8Anthropicclaude-sonnet-4-61472.0 to 1483.6-35.118,185
131477.2OpenAIgpt-5.5-high1470.7 to 1483.7-35.714,997
141475.7OpenAIgpt-5.4-high1469.9 to 1481.4-37.219,104
151474.8Anthropicclaude-opus-4-5-202511011469.6 to 1479.9-38.120,674
161474.0GoogleGemini 3 Pro1467.6 to 1480.4-38.911,180
171470.5OpenAIgpt-5.51464.1 to 1476.8-42.415,625
181470.2Xiaomimimo-v2.5-pro1463.6 to 1476.7-42.713,488
191469.8Alibabaqwen3.5-max-preview1462.1 to 1477.5-43.16,866
201468.3Alibabaqwen3.7-max-preview1451.1 to 1485.4-44.61,281
211464.3Anthropicclaude-sonnet-4-5-20250929-thinking-32k1459.6 to 1468.9-48.623,579
221464.0Z.aiglm-5.2 (max)1455.5 to 1472.5-48.95,859
231463.7Googlegemini-3.5-flash-medium1454.3 to 1473.1-49.24,586
241463.6Z.aiglm-5.11456.9 to 1470.4-49.310,098
251462.9Anthropicclaude-sonnet-4-5-202509291458.2 to 1467.6-50.023,247

Frontier race

Category Kings: Instruction following arena frontier by organization

Anthropic holds the instruction following arena frontier at 1513 as of July 20, 2026.

Instruction following arena frontier ratings by organization over timeAnthropic holds the instruction following arena frontier at 1513 as of July 20, 2026. 10 organizations are tracked from January 5, 2025 to July 20, 2026.12101300139014802025-012025-052025-092026-012026-042026-07Anthropic 1513OpenAI 1484Google 1481Meta 1478Xiaomi 1470Alibaba 1470Z.ai 1464xAI 1460DeepSeek 1454Mistral AI 1421

Method: Frontier is the maximum resolved LMArena rating per organization on each publish date. Organizations limited to the union of the seven tracked homepage providers and organizations in the latest board top 25 by officialRank. Quarantined dates are excluded, and this category series keeps its own method epoch and is never joined to the legacy archive.

Instruction following arena frontier by organization: latest resolved score per organization, from the reviewed LMArena category artifact.
OrganizationLatest modelLatest LMArena ratingFirst tracked date
Anthropicanthropic/claude-fable-51512.9January 5, 2025
OpenAIopenai/gpt-5.6-sol-xhigh1483.5January 5, 2025
Googlegoogle/gemini-3.1-pro-preview1481.2January 5, 2025
Metameta/muse-spark-1.11478.1January 5, 2025
Xiaomixiaomi/mimo-v2.5-pro1470.2December 23, 2025
Alibabaalibaba/qwen3.5-max-preview1469.8January 5, 2025
Z.aizai/glm-5.2-max1464.0January 5, 2025
xAIxai/grok-4.51460.0January 5, 2025
DeepSeekdeepseek/deepseek-v4-pro1453.5January 5, 2025
Mistral AImistral/mistral-medium-3.51421.1January 5, 2025
Anthropic leads the instruction following arena frontier at 1512.9 on July 20, 2026. Frontier is the highest resolved score per organization per publish date. Quarantined dates are excluded and this category is never joined to the legacy archive.

Definitions

What these terms mean.

Instruction following arena
The LMArena style-controlled text battles filtered to prompts categorized as instruction following. It measures preference on constraint adherence, not verified compliance rates.
Frontier by organization
The highest resolved score for each organization on each publish date. It tracks how far each lab has pushed this category, not an average of all its models.
Confidence interval
The reported 95% range LMArena publishes with each rating. Overlapping intervals mean this snapshot does not separate those ranks decisively.

Instruction following arena questions

Which model leads the Instruction following arena right now?

According to the LMArena instruction following board published July 20, 2026, Anthropic leads with claude-fable-5 at 1512.9, across 3,433 battles. That is 0.1 points ahead of the next board entry inside this one snapshot.

Is this a Model Gauntlet instruction following ranking?

No. This board mirrors one named source, the LMArena instruction following arena, published July 20, 2026. Model Gauntlet publishes no composite or consensus score and runs no evaluation of its own.

How many models and organizations are in the instruction following board?

The reviewed artifact covers 382 models across 61 organizations, with 378 models resolved to exact identities and 4 held for review. The history spans 145 snapshots from January 5, 2025 to July 20, 2026.

Why do some scores have overlapping confidence intervals?

LMArena publishes a 95% confidence interval with each rating. Where two intervals overlap, this single snapshot does not separate those ranks decisively, so treat small instruction following gaps with caution.