← Overall rankings

LMArena Math arena · published July 20, 2026

Anthropic leads the LMArena math arena.

As of July 20, 2026, Anthropic leads the LMArena math arena with claude-fable-5 at 1550.2 across 487 battles, according to LMArena.

Current through July 20, 2026. 145 published snapshots. 373 of 376 models resolved (3 held for exact-match review).

LMArena math arena rating, published July 20, 2026. Top 25 of 316 rows.
RankLMArena ratingOrganizationModelReported intervalGap to leaderBattles
011550.2AnthropicClaude Fable 51524.0 to 1576.4Leader487
021518.8Anthropicclaude-opus-4-6-thinking1507.8 to 1529.9-31.43,231
031517.6Googlegemini-3.5-flash-high1492.2 to 1543.0-32.6572
041504.5Anthropicclaude-opus-4-7-thinking1492.2 to 1516.8-45.72,532
051503.5xAIgrok-4.51469.5 to 1537.6-46.7287
061502.1Anthropicclaude-opus-4-61491.6 to 1512.5-48.13,651
071496.7OpenAIgpt-5.51483.9 to 1509.5-53.52,376
081496.3Anthropicclaude-opus-4-8-thinking1480.8 to 1511.8-53.91,471
091494.9OpenAIgpt-5.4-high1483.6 to 1506.1-55.33,042
101493.4AnthropicClaude Opus 4.71481.2 to 1505.6-56.82,627
111491.3Alibabaqwen3.7-max-preview1451.3 to 1531.4-58.9219
121490.9Googlegemini-3.1-pro-preview1481.3 to 1500.4-59.34,502
131489.9OpenAIgpt-5.5-high1477.4 to 1502.5-60.32,328
141481.1Metamuse-spark-1.11448.5 to 1513.7-69.1320
151480.5Z.aiglm-5.11465.2 to 1495.9-69.71,559
161479.8Moonshot AIkimi-k2.61466.2 to 1493.4-70.41,931
171479.5Googlegemini-3.5-flash-medium1456.6 to 1502.5-70.7621
181478.3Baiduernie-5.11464.5 to 1492.0-71.91,899
191478.0Anthropicclaude-opus-4-81462.6 to 1493.5-72.21,487
201477.7GoogleGemini 3 Pro1466.2 to 1489.1-72.52,645
211476.8Alibabaqwen3.6-max-preview1446.8 to 1506.7-73.4359
221476.3Googlegemini-3-flash1463.2 to 1489.5-73.91,991
231475.3Xiaomimimo-v2.5-pro1461.8 to 1488.8-74.92,036
241475.0Z.aiglm-5.2 (max)1454.9 to 1495.2-75.2829
251474.0Alibabaqwen3.7-plus1455.3 to 1492.7-76.2996

Frontier race

Category Kings: Math arena frontier by organization

Anthropic holds the math arena frontier at 1550 as of July 20, 2026.

Math arena frontier ratings by organization over timeAnthropic holds the math arena frontier at 1550 as of July 20, 2026. 12 organizations are tracked from January 5, 2025 to July 20, 2026.11801290140015102025-012025-052025-092026-012026-042026-07Anthropic 1550Google 1518xAI 1504OpenAI 1497Alibaba 1491Meta 1481Z.ai 1481Moonshot AI 1480Baidu 1478Xiaomi 1475DeepSeek 1470Mistral AI 1433

Method: Frontier is the maximum resolved LMArena rating per organization on each publish date. Organizations limited to the union of the seven tracked homepage providers and organizations in the latest board top 25 by officialRank. Quarantined dates are excluded, and this category series keeps its own method epoch and is never joined to the legacy archive.

Math arena frontier by organization: latest resolved score per organization, from the reviewed LMArena category artifact.
OrganizationLatest modelLatest LMArena ratingFirst tracked date
Anthropicanthropic/claude-fable-51550.2January 5, 2025
Googlegoogle/gemini-3.5-flash-high1517.6January 5, 2025
xAIxai/grok-4.51503.5January 5, 2025
OpenAIopenai/gpt-5.51496.7January 5, 2025
Alibabaalibaba/qwen3.7-max-preview1491.3January 5, 2025
Metameta/muse-spark-1.11481.1January 5, 2025
Z.aizai/glm-5.11480.5January 5, 2025
Moonshot AImoonshot/kimi-k2.61479.8July 17, 2025
Baidubaidu/ernie-5.11478.3November 16, 2025
Xiaomixiaomi/mimo-v2.5-pro1475.3December 30, 2025
DeepSeekdeepseek/deepseek-v4-pro-thinking1469.8January 5, 2025
Mistral AImistral/mistral-medium-3.51432.6January 5, 2025
Anthropic leads the math arena frontier at 1550.2 on July 20, 2026. Frontier is the highest resolved score per organization per publish date. Quarantined dates are excluded and this category is never joined to the legacy archive.

Definitions

What these terms mean.

Math arena
The LMArena style-controlled text battles filtered to prompts categorized as math. It reflects human preference on math prompts, not verified proof correctness.
Frontier by organization
The highest resolved score for each organization on each publish date. It tracks how far each lab has pushed this category, not an average of all its models.
Confidence interval
The reported 95% range LMArena publishes with each rating. Overlapping intervals mean this snapshot does not separate those ranks decisively.

Math arena questions

Which model leads the Math arena right now?

According to the LMArena math board published July 20, 2026, Anthropic leads with claude-fable-5 at 1550.2, across 487 battles. That is 31.4 points ahead of the next board entry inside this one snapshot.

Is this a Model Gauntlet math ranking?

No. This board mirrors one named source, the LMArena math arena, published July 20, 2026. Model Gauntlet publishes no composite or consensus score and runs no evaluation of its own.

How many models and organizations are in the math board?

The reviewed artifact covers 376 models across 60 organizations, with 373 models resolved to exact identities and 3 held for review. The history spans 145 snapshots from January 5, 2025 to July 20, 2026.

Why do some scores have overlapping confidence intervals?

LMArena publishes a 95% confidence interval with each rating. Where two intervals overlap, this single snapshot does not separate those ranks decisively, so treat small math gaps with caution.