← Overall rankings

LMArena Long queries arena · published July 20, 2026

Anthropic leads the LMArena long queries arena.

As of July 20, 2026, Anthropic leads the LMArena long queries arena with claude-opus-4-6-thinking at 1524.0 across 24,855 battles, according to LMArena.

Current through July 20, 2026. 145 published snapshots. 378 of 382 models resolved (4 held for exact-match review).

LMArena long queries arena rating, published July 20, 2026. Top 25 of 319 rows.
RankLMArena ratingOrganizationModelReported intervalGap to leaderBattles
011524.0Anthropicclaude-opus-4-6-thinking1518.5 to 1529.5Leader24,855
021524.0AnthropicClaude Fable 51514.1 to 1533.8-0.04,362
031517.0Anthropicclaude-opus-4-7-thinking1511.1 to 1523.0-7.021,959
041517.0Anthropicclaude-opus-4-61511.7 to 1522.4-7.027,072
051508.8AnthropicClaude Opus 4.71502.9 to 1514.7-15.222,653
061506.2Anthropicclaude-opus-4-8-thinking1499.2 to 1513.2-17.813,426
071500.4Anthropicclaude-opus-4-81493.4 to 1507.5-23.613,899
081497.8Alibabaqwen3.7-max-preview1482.0 to 1513.6-26.21,597
091497.7Googlegemini-3.1-pro-preview1492.6 to 1502.8-26.334,336
101494.8Anthropicclaude-opus-4-5-20251101-thinking-32k1488.0 to 1501.6-29.29,235
111494.3Anthropicclaude-sonnet-4-61488.7 to 1499.9-29.723,226
121492.4Anthropicclaude-opus-4-5-202511011487.3 to 1497.6-31.622,656
131492.4GoogleGemini 3 Pro1485.8 to 1499.0-31.610,480
141489.1OpenAIgpt-5.6-sol-xhigh1476.1 to 1502.1-34.92,324
151487.1Xiaomimimo-v2.5-pro1480.8 to 1493.4-36.917,664
161485.5Anthropicclaude-sonnet-4-5-20250929-thinking-32k1480.8 to 1490.2-38.524,489
171485.0OpenAIgpt-5.5-high1478.9 to 1491.1-39.019,765
181484.9Anthropicclaude-opus-4-1-20250805-thinking-16k1478.7 to 1491.1-39.111,128
191484.1Z.aiglm-5.11477.7 to 1490.6-39.913,101
201484.0Metamuse-spark-1.11472.9 to 1495.1-40.03,145
211483.9OpenAIgpt-5.4-high1478.3 to 1489.5-40.124,527
221483.4Anthropicclaude-sonnet-4-5-202509291478.6 to 1488.2-40.624,885
231482.0xAIgrok-4.51470.8 to 1493.2-42.03,138
241481.1Alibabaqwen3.5-max-preview1473.7 to 1488.5-42.98,259
251481.0Z.aiglm-5.2 (max)1473.0 to 1489.0-43.07,693

Frontier race

Category Kings: Long queries arena frontier by organization

Anthropic holds the long queries arena frontier at 1524 as of July 20, 2026.

Long queries arena frontier ratings by organization over timeAnthropic holds the long queries arena frontier at 1524 as of July 20, 2026. 10 organizations are tracked from January 5, 2025 to July 20, 2026.121013001390148015702025-012025-052025-092026-012026-042026-07Anthropic 1524Alibaba 1498Google 1498OpenAI 1489Xiaomi 1487Z.ai 1484Meta 1484xAI 1482DeepSeek 1473Mistral AI 1431

Method: Frontier is the maximum resolved LMArena rating per organization on each publish date. Organizations limited to the union of the seven tracked homepage providers and organizations in the latest board top 25 by officialRank. Quarantined dates are excluded, and this category series keeps its own method epoch and is never joined to the legacy archive.

Long queries arena frontier by organization: latest resolved score per organization, from the reviewed LMArena category artifact.
OrganizationLatest modelLatest LMArena ratingFirst tracked date
Anthropicanthropic/claude-fable-51524.0January 5, 2025
Alibabaalibaba/qwen3.7-max-preview1497.8January 5, 2025
Googlegoogle/gemini-3.1-pro-preview1497.7January 5, 2025
OpenAIopenai/gpt-5.6-sol-xhigh1489.1January 5, 2025
Xiaomixiaomi/mimo-v2.5-pro1487.1December 23, 2025
Z.aizai/glm-5.11484.1January 5, 2025
Metameta/muse-spark-1.11484.0January 5, 2025
xAIxai/grok-4.51482.0January 5, 2025
DeepSeekdeepseek/deepseek-v4-pro1473.2January 5, 2025
Mistral AImistral/mistral-medium-3.51431.1January 5, 2025
Anthropic leads the long queries arena frontier at 1524.0 on July 20, 2026. Frontier is the highest resolved score per organization per publish date. Quarantined dates are excluded and this category is never joined to the legacy archive.

Definitions

What these terms mean.

Long queries arena
The LMArena style-controlled text battles filtered to longer prompts as classified by the source. Prompt length is the source's cutoff, not a context-window benchmark.
Frontier by organization
The highest resolved score for each organization on each publish date. It tracks how far each lab has pushed this category, not an average of all its models.
Confidence interval
The reported 95% range LMArena publishes with each rating. Overlapping intervals mean this snapshot does not separate those ranks decisively.

Long queries arena questions

Which model leads the Long queries arena right now?

According to the LMArena long queries board published July 20, 2026, Anthropic leads with claude-opus-4-6-thinking at 1524.0, across 24,855 battles. That is 0.0 points ahead of the next board entry inside this one snapshot.

Is this a Model Gauntlet long queries ranking?

No. This board mirrors one named source, the LMArena long queries arena, published July 20, 2026. Model Gauntlet publishes no composite or consensus score and runs no evaluation of its own.

How many models and organizations are in the long queries board?

The reviewed artifact covers 382 models across 61 organizations, with 378 models resolved to exact identities and 4 held for review. The history spans 145 snapshots from January 5, 2025 to July 20, 2026.

Why do some scores have overlapping confidence intervals?

LMArena publishes a 95% confidence interval with each rating. Where two intervals overlap, this single snapshot does not separate those ranks decisively, so treat small long queries gaps with caution.