← Overall rankings

Agent Arena · published July 20, 2026

Claude Fable 5 (High) leads the agent board.

LMArena's Agent Arena scores models on real agentic sessions and published this 37-row board July 20, 2026. IPS is the arena's own metric on its own scale. Configuration-labeled rows are system results, and nothing here is compared against any other board.

Agent Arena overall standings published July 20, 2026.
RankSystemOrganizationIPS scoreIdentity
01Claude Fable 5 (High)anthropic0.132reviewed
02Claude Opus 4.8 (Thinking)anthropic0.100reviewed
03GPT 5.6 Sol (xHigh)openai0.099reviewed
04Kimi K3moonshot0.096held
05Claude Sonnet 5 (High)anthropic0.091reviewed
06GPT 5.5 (xHigh)openai0.088reviewed
07Claude Opus 4.7 (Thinking)anthropic0.083reviewed
08Claude Opus 4.7anthropic0.081reviewed
09GPT 5.5 (High)openai0.077reviewed
10Claude Opus 4.6anthropic0.067reviewed
11GLM 5.2 (Max)zai0.065reviewed
12GPT 5.5openai0.064reviewed
13Grok 4.5xai0.064reviewed
14GPT 5.4 (High)openai0.059reviewed
15Claude Opus 4.8anthropic0.040reviewed
16Claude Sonnet 4.6anthropic0.032reviewed
17GLM 5.1zai0.017reviewed
18Muse Spark 1.1meta0.012reviewed
19Qwen3.7 Maxalibaba0.003reviewed
20Gemini 3.1 Pro Previewgoogle-0.002reviewed
21Gemini 3.5 Flash (High)google-0.007reviewed
22Kimi K2.7 Codemoonshot-0.007reviewed
23Qwen3.7 Plusalibaba-0.008reviewed
24DeepSeek V4 Prodeepseek-0.008reviewed
25Kimi K2.6moonshot-0.023reviewed
26Minimax M3minimax-0.027reviewed
27Mimo V2.5 Proxiaomi-0.030reviewed
28DeepSeek V4 Flashdeepseek-0.038reviewed
29Inklingthinky-0.057held
30Gemini 3.5 Flash (Medium)google-0.065reviewed
31Grok Build 0.1xai-0.076held
32Grok 4.3 (High)xai-0.078reviewed
33Gemini 3 Flashgoogle-0.083reviewed
34Minimax M2.7minimax-0.115reviewed
35Nemotron 3 Ultranvidia-0.129reviewed
36Gemma 4 31Bgoogle-0.140reviewed
37Grok 4.3xai-0.147reviewed

Agent board questions

What does the Agent Arena measure?

LMArena's Agent Arena evaluates models on real agentic sessions, where a model plans and executes multi-step work, and scores them with IPS, the arena's own metric. Higher is better, and the rank ladder is the arena's own published order.

Why do some rows name a configuration like (High) or (Thinking)?

Those rows are system results for a specific reasoning configuration, not bare-model results. The same base model can appear more than once at different configurations, and each row keeps the arena's own label.

Is an IPS score comparable to an LMArena rating?

No. IPS is this arena's own scale, and this site never merges, compares, or rescales it against any other board. The board shown here was published July 20, 2026 inside LMArena's CC BY 4.0 leaderboard dataset.