Agent Arena · published July 20, 2026
Claude Fable 5 (High) leads the agent board.
LMArena's Agent Arena scores models on real agentic sessions and published this 37-row board July 20, 2026. IPS is the arena's own metric on its own scale. Configuration-labeled rows are system results, and nothing here is compared against any other board.
| Rank | System | Organization | IPS score | Identity |
|---|---|---|---|---|
| 01 | Claude Fable 5 (High) | anthropic | 0.132 | reviewed |
| 02 | Claude Opus 4.8 (Thinking) | anthropic | 0.100 | reviewed |
| 03 | GPT 5.6 Sol (xHigh) | openai | 0.099 | reviewed |
| 04 | Kimi K3 | moonshot | 0.096 | held |
| 05 | Claude Sonnet 5 (High) | anthropic | 0.091 | reviewed |
| 06 | GPT 5.5 (xHigh) | openai | 0.088 | reviewed |
| 07 | Claude Opus 4.7 (Thinking) | anthropic | 0.083 | reviewed |
| 08 | Claude Opus 4.7 | anthropic | 0.081 | reviewed |
| 09 | GPT 5.5 (High) | openai | 0.077 | reviewed |
| 10 | Claude Opus 4.6 | anthropic | 0.067 | reviewed |
| 11 | GLM 5.2 (Max) | zai | 0.065 | reviewed |
| 12 | GPT 5.5 | openai | 0.064 | reviewed |
| 13 | Grok 4.5 | xai | 0.064 | reviewed |
| 14 | GPT 5.4 (High) | openai | 0.059 | reviewed |
| 15 | Claude Opus 4.8 | anthropic | 0.040 | reviewed |
| 16 | Claude Sonnet 4.6 | anthropic | 0.032 | reviewed |
| 17 | GLM 5.1 | zai | 0.017 | reviewed |
| 18 | Muse Spark 1.1 | meta | 0.012 | reviewed |
| 19 | Qwen3.7 Max | alibaba | 0.003 | reviewed |
| 20 | Gemini 3.1 Pro Preview | -0.002 | reviewed | |
| 21 | Gemini 3.5 Flash (High) | -0.007 | reviewed | |
| 22 | Kimi K2.7 Code | moonshot | -0.007 | reviewed |
| 23 | Qwen3.7 Plus | alibaba | -0.008 | reviewed |
| 24 | DeepSeek V4 Pro | deepseek | -0.008 | reviewed |
| 25 | Kimi K2.6 | moonshot | -0.023 | reviewed |
| 26 | Minimax M3 | minimax | -0.027 | reviewed |
| 27 | Mimo V2.5 Pro | xiaomi | -0.030 | reviewed |
| 28 | DeepSeek V4 Flash | deepseek | -0.038 | reviewed |
| 29 | Inkling | thinky | -0.057 | held |
| 30 | Gemini 3.5 Flash (Medium) | -0.065 | reviewed | |
| 31 | Grok Build 0.1 | xai | -0.076 | held |
| 32 | Grok 4.3 (High) | xai | -0.078 | reviewed |
| 33 | Gemini 3 Flash | -0.083 | reviewed | |
| 34 | Minimax M2.7 | minimax | -0.115 | reviewed |
| 35 | Nemotron 3 Ultra | nvidia | -0.129 | reviewed |
| 36 | Gemma 4 31B | -0.140 | reviewed | |
| 37 | Grok 4.3 | xai | -0.147 | reviewed |
Agent board questions
What does the Agent Arena measure?
LMArena's Agent Arena evaluates models on real agentic sessions, where a model plans and executes multi-step work, and scores them with IPS, the arena's own metric. Higher is better, and the rank ladder is the arena's own published order.
Why do some rows name a configuration like (High) or (Thinking)?
Those rows are system results for a specific reasoning configuration, not bare-model results. The same base model can appear more than once at different configurations, and each row keeps the arena's own label.
Is an IPS score comparable to an LMArena rating?
No. IPS is this arena's own scale, and this site never merges, compares, or rescales it against any other board. The board shown here was published July 20, 2026 inside LMArena's CC BY 4.0 leaderboard dataset.