LMArena Web development arena · published May 4, 2026
Anthropic leads the LMArena web development arena.
As of May 4, 2026, Anthropic leads the LMArena web development arena with claude-opus-4-7-thinking at 1567.9 across 2,948 battles, according to LMArena.
Current through May 4, 2026. 44 published snapshots. 56 of 72 models resolved (16 held for exact-match review).
| Rank | LMArena rating | Organization | Model | Reported interval | Gap to leader | Battles |
|---|---|---|---|---|---|---|
| 01 | 1567.9 | Anthropic | claude-opus-4-7-thinking | 1555.4 to 1580.3 | Leader | 2,948 |
| 02 | 1561.9 | Anthropic | Claude Opus 4.7 | 1549.8 to 1574.0 | -6.0 | 2,999 |
| 03 | 1548.8 | Anthropic | claude-opus-4-6-thinking | 1539.9 to 1557.8 | -19.1 | 6,075 |
| 04 | 1544.4 | Anthropic | claude-opus-4-6 | 1535.9 to 1552.8 | -23.5 | 7,030 |
| 05 | 1531.7 | Z.ai | glm-5.1 | 1520.7 to 1542.7 | -36.2 | 3,602 |
| 06 | 1526.2 | Anthropic | claude-sonnet-4-6 | 1518.4 to 1533.9 | -41.7 | 9,084 |
| 07 | 1524.6 | Moonshot AI | kimi-k2.6 | 1511.6 to 1537.6 | -43.3 | 2,392 |
| 08 | 1509.4 | Meta | Muse Spark | 1493.5 to 1525.3 | -58.5 | 1,634 |
| 09 | 1490.6 | Anthropic | claude-opus-4-5-20251101-thinking-32k | 1483.4 to 1497.8 | -77.3 | 13,060 |
| 10 | 1476.4 | Xiaomi | mimo-v2.5-pro | 1464.7 to 1488.2 | -91.5 | 2,992 |
| 11 | 1467.2 | Anthropic | claude-opus-4-5-20251101 | 1460.8 to 1473.6 | -100.7 | 15,298 |
| 12 | 1464.0 | Alibaba | qwen3.6-plus | 1454.0 to 1474.0 | -103.9 | 4,099 |
| 13 | 1454.7 | gemini-3.1-pro-preview | 1447.3 to 1462.2 | -113.2 | 8,208 | |
| 14 | 1454.7 | DeepSeek | deepseek-v4-pro-thinking | 1438.5 to 1470.8 | -113.2 | 1,470 |
| 15 | 1445.6 | Xiaomi | mimo-v2.5 | 1431.4 to 1459.7 | -122.3 | 1,891 |
| 16 | 1439.9 | Z.ai | glm-4.7 | 1429.6 to 1450.2 | -128.0 | 4,883 |
| 17 | 1438.2 | Gemini 3 Pro | 1430.9 to 1445.5 | -129.7 | 17,163 | |
| 18 | 1437.0 | gemini-3-flash | 1429.6 to 1444.5 | -130.9 | 13,281 | |
| 19 | 1435.9 | Z.ai | glm-5 | 1427.7 to 1444.1 | -132.0 | 6,567 |
| 20 | 1429.7 | Moonshot AI | kimi-k2.5-thinking | 1422.5 to 1436.9 | -138.2 | 9,006 |
| 21 | 1428.5 | Xiaomi | mimo-v2-pro | 1419.5 to 1437.6 | -139.4 | 5,127 |
| 22 | 1408.7 | MiniMax | minimax-m2.7 | 1399.4 to 1418.0 | -159.2 | 4,750 |
| 23 | 1407.9 | Moonshot AI | kimi-k2.5-instant | 1397.2 to 1418.5 | -160.0 | 3,610 |
| 24 | 1405.0 | xAI | grok-4.3 | 1389.8 to 1420.2 | -162.9 | 1,585 |
| 25 | 1403.6 | OpenAI | gpt-5.2 | 1387.0 to 1420.3 | -164.3 | 1,457 |
Frontier race
Category Kings: Web development arena frontier by organization
Anthropic holds the web development arena frontier at 1568 as of May 4, 2026.
Method: Frontier is the maximum resolved LMArena rating per organization on each publish date. Organizations limited to the union of the seven tracked homepage providers and organizations in the latest board top 25 by officialRank. Quarantined dates are excluded, and this category series keeps its own method epoch and is never joined to the legacy archive.
| Organization | Latest model | Latest LMArena rating | First tracked date |
|---|---|---|---|
| Anthropic | anthropic/claude-opus-4-7-thinking | 1567.9 | November 12, 2025 |
| Z.ai | zai/glm-5.1 | 1531.7 | November 12, 2025 |
| Moonshot AI | moonshot/kimi-k2.6 | 1524.6 | November 16, 2025 |
| Meta | meta/muse-spark | 1509.4 | April 22, 2026 |
| Xiaomi | xiaomi/mimo-v2.5-pro | 1476.4 | December 23, 2025 |
| Alibaba | alibaba/qwen3.6-plus | 1464.0 | November 16, 2025 |
| DeepSeek | deepseek/deepseek-v4-pro-thinking | 1454.7 | November 16, 2025 |
| google/gemini-3.1-pro-preview | 1454.7 | November 12, 2025 | |
| MiniMax | minimax/minimax-m2.7 | 1408.7 | November 12, 2025 |
| xAI | xai/grok-4.3 | 1405.0 | November 25, 2025 |
| OpenAI | openai/gpt-5.2 | 1403.6 | November 16, 2025 |
| Mistral AI | mistral/mistral-large-3 | 1222.4 | December 11, 2025 |
Definitions
What these terms mean.
- Web development arena
- The LMArena web development arena, where models build web interfaces judged by human preference. LMArena stores this board under the webdev label.
- Frontier by organization
- The highest resolved score for each organization on each publish date. It tracks how far each lab has pushed this category, not an average of all its models.
- Confidence interval
- The reported 95% range LMArena publishes with each rating. Overlapping intervals mean this snapshot does not separate those ranks decisively.
Follow the evidence
Compare this board against the overall LMArena rankings, shortlist candidates with the Coding shortlist guide,read the Anthropic lab file, and see how rankings are selected. Model profiles on this board: Claude Opus 4.7, Muse Spark, Gemini 3 Pro.
Web development arena questions
Which model leads the Web development arena right now?
According to the LMArena web development board published May 4, 2026, Anthropic leads with claude-opus-4-7-thinking at 1567.9, across 2,948 battles. That is 6.0 points ahead of the next board entry inside this one snapshot.
Is this a Model Gauntlet web development ranking?
No. This board mirrors one named source, the LMArena web development arena, published May 4, 2026. Model Gauntlet publishes no composite or consensus score and runs no evaluation of its own.
How many models and organizations are in the web development board?
The reviewed artifact covers 72 models across 15 organizations, with 56 models resolved to exact identities and 16 held for review. The history spans 44 snapshots from November 12, 2025 to May 4, 2026.
Why do some scores have overlapping confidence intervals?
LMArena publishes a 95% confidence interval with each rating. Where two intervals overlap, this single snapshot does not separate those ranks decisively, so treat small web development gaps with caution.