MODEL GAUNTLET RESEARCH DESK

CURRENT RANKING SNAPSHOTUPDATED JULY 16, 2026

SOURCE-ATTRIBUTED AI RANKING AGGREGATOR EDITORIALLY INDEPENDENT

Who’s winningthe AI race?

Anthropic leads the July 16, 2026 LMArena text ranking.

MODELGAUNTLET.COMONE SOURCE VIEW · REPORTED UNCERTAINTY SHOWNANTHROPIC +14.4 OVER META · LMARENA JULY 16, 2026

The current frontier

The models in the race right now.

An editorial watchlist of the models in active frontier discussion, audited July 22, 2026: every pick appears on at least one current reviewed board. Selection is ours; every number on a card is one named source's own value with its own scale, every card carries the same stat rows, and a board that does not list a model says so. Nothing is blended.

LMArena rating · style-controlled text · published July 20, 2026

Bars start at zero on LMArena's own scale, so close ratings look close. Hover or focus a bar for the reported interval and battle count. The cards below carry the same values as text.

Claude Fable 5

AnthropicReleased June 9, 2026

LMArena rank
#1 · 1506.8
Agent Arena
#1 · Claude Fable 5 (High)
OpenRouter usage
#24 · 449.9B tokens/week
OpenRouter price
$10.00 in / $50.00 out per 1M

Claude Mythos 5

AnthropicReleased June 9, 2026

LMArena rank
Not on this board
Agent Arena
Not on this board
OpenRouter usage
No matched listing
OpenRouter price
Not listed

Claude Opus 4.8 Thinking

AnthropicNot in the reviewed catalog

LMArena rank
#12 · 1484.6
Agent Arena
#2 · Claude Opus 4.8 (Thinking)
OpenRouter usage
No matched listing
OpenRouter price
Not listed

Claude Sonnet 5 High

AnthropicNot in the reviewed catalog

LMArena rank
#37 · 1461.4
Agent Arena
#5 · Claude Sonnet 5 (High)
OpenRouter usage
No matched listing
OpenRouter price
Not listed

GPT-5.6 Sol X-High

OpenAINot in the reviewed catalog

LMArena rank
#9 · 1486.4
Agent Arena
#3 · GPT 5.6 Sol (xHigh)
OpenRouter usage
No matched listing
OpenRouter price
Not listed

GPT-5.5 High

OpenAINot in the reviewed catalog

LMArena rank
#13 · 1481.0
Agent Arena
#9 · GPT 5.5 (High)
OpenRouter usage
No matched listing
OpenRouter price
Not listed

Gemini 3.5 Flash High

GoogleNot in the reviewed catalog

LMArena rank
#15 · 1476.3
Agent Arena
#21 · Gemini 3.5 Flash (High)
OpenRouter usage
No matched listing
OpenRouter price
Not listed

Gemini 3.1 Pro Preview

GoogleNot in the reviewed catalog

LMArena rank
#11 · 1485.6
Agent Arena
#20 · Gemini 3.1 Pro Preview
OpenRouter usage
#37 · 174.6B tokens/week
OpenRouter price
$2.00 in / $12.00 out per 1M
EQ-Bench 3
#19 · 1302

7 Epoch hub benchmarks carry this model. See the benchmark section.

Grok 4.5

xAINot in the reviewed catalog

LMArena rank
#35 · 1465.7
Agent Arena
#13 · Grok 4.5
OpenRouter usage
#28 · 355.1B tokens/week
OpenRouter price
$2.00 in / $6.00 out per 1M

DeepSeek V4 Pro

DeepSeekReleased April 24, 2026

LMArena rank
#43 · 1457.1
Agent Arena
#24 · DeepSeek V4 Pro
OpenRouter usage
#6 · 2.8T tokens/week
OpenRouter price
$0.43 in / $0.87 out per 1M
EQ-Bench 3
#5 · 1330

GLM 5.2 (Max)

Z.aiNot in the reviewed catalog

LMArena rank
#29 · 1469.5
Agent Arena
#11 · GLM 5.2 (Max)
OpenRouter usage
No matched listing
OpenRouter price
Not listed

Kimi K3

Moonshot AINot in the reviewed catalog

LMArena rank
Not on this board
Agent Arena
#4 · Kimi K3
OpenRouter usage
No matched listing
OpenRouter price
Not listed

Qwen3.7 Max Preview

AlibabaNot in the reviewed catalog

LMArena rank
#18 · 1475.3
Agent Arena
Not on this board
OpenRouter usage
No matched listing
OpenRouter price
Not listed

MiMo V2.5

XiaomiNot in the reviewed catalog

LMArena rank
#78 · 1433.1
Agent Arena
Not on this board
OpenRouter usage
#2 · 9.4T tokens/week
OpenRouter price
Not listed

Sources: LMArena category artifact published July 20, 2026, Aider polyglot updated May 22, 2026, EQ-Bench 3 updated May 10, 2026, and the Epoch AI Benchmarking Hub retrieved July 19, 2026. Each carries its receipt in its labeled section below. Rows named with a scaffold are system results, not bare-model results.

Chatbot Arena archive

Sampled provider frontier ratings, May 2023 to August 2025

Meta highlighted. 827 to 1288 published rating; sampled peak 1396.

Sampled provider frontier ratings over time Google leads the final sampled snapshot at 1466.2 on August 29, 2025. Meta moved from 827.0 to 1288.0 across the selected snapshots and peaked at 1395.9.8001000120014002023-052023-102024-032024-082025-022025-082025-08-29SNAPSHOTGoogle 1466Mistral 1433xAI 1431OpenAI 1429DeepSeek 1424Anthropic 1417Meta 1288

Line ends, top to bottom:Google 1466Mistral 1433xAI 1431OpenAI 1429DeepSeek 1424Anthropic 1417Meta 1288

Eleven selected published snapshots are shown. The selected snapshots cross LMArena rating-method changes, including Bradley-Terry, style control, and frequency reweighting. Endpoint differences are not normalized like-for-like performance changes. This legacy series is never joined to the current table. Each raw snapshot has a published SHA-256 receipt.Why the chart stops in August 2025 →

Current LMArena snapshot

See who leads this source, and how close the race really is.

These values come from LMArena’s text style-control overall table published July 16, 2026. The legacy archive remains separate.

1507.5

Anthropic leads

Claude Fable 5 is the highest-scoring model in LMArena’s July 16, 2026 snapshot, with a reported interval of 1500.2 to 1514.7.

14.4

Gap to second

Meta’s Muse Spark 1.1 ranks second at 1493.1, 14.4 points behind Anthropic.

8,817

Leader battles

LMArena reports 8,817 battles for Claude Fable 5 in this published snapshot.

7

Tracked organizations

Each model is matched to the organization listed in LMArena’s published data, avoiding guesses based on model names.

Point-in-time standings

The leading model from each AI lab in this source.

We take each organization’s highest LMArena rating in this published snapshot. Confidence intervals show where close results may not represent a meaningful difference.

LMArena rating, style-controlled text category, published July 16, 2026. Higher is better.
Source rankModelOrganizationLMArena ratingReported intervalBattles
01Claude Fable 5Anthropic1507.51500.2–1514.78,817
02Muse Spark 1.1Meta1493.11484.9–1501.35,732
03GPT-5.6 Sol X-HighOpenAI1486.31476.8–1495.84,113
04Gemini 3 ProGoogle1485.81481.9–1489.641,283
05Grok 4.20 Beta 1xAI1474.21469.5–1478.826,844
06DeepSeek V4 ProDeepSeek1457.51453.1–1462.043,062
07Mistral Medium 3.5Mistral1427.31420.7–1433.911,017

LMArena category boards

Six task boards from the same named source.

Each board is LMArena's own category rating. Open one board at a time; every board links to its full page with intervals and history.

Coding arenaTop 10 of 322 resolved rows
  1. 01Claude Opus 4.7 Thinking1553.0
  2. 02Claude Opus 4.6 Thinking1550.1
  3. 03Claude Opus 4.71548.7
  4. 04Claude Fable 51547.7
  5. 05Claude Opus 4.61547.6
  6. 06Claude Opus 4.8 Thinking1535.1
  7. 07Muse Spark 1.11533.0
  8. 08Claude Opus 4.81530.7
  9. 09Claude Opus 4.5 Thinking 32K (2025-11-01)1530.3
  10. 10Claude Sonnet 4.61527.4
Open the full coding arena board →
Creative writing arenaTop 10 of 322 resolved rows
  1. 01Claude Fable 51512.8
  2. 02Claude Opus 4.6 Thinking1499.1
  3. 03Claude Opus 4.7 Thinking1487.8
  4. 04Gemini 3 Pro1484.3
  5. 05Claude Opus 4.71482.6
  6. 06GPT-5.6 Sol X-High1481.2
  7. 07Claude Opus 4.61479.3
  8. 08Gemini 3.1 Pro Preview1477.9
  9. 09Claude Opus 4.5 Thinking 32K (2025-11-01)1468.0
  10. 10Gemini 3.5 Flash Medium1466.8
Open the full creative writing arena board →
Instruction following arenaTop 10 of 324 resolved rows
  1. 01Claude Fable 51512.9
  2. 02Claude Opus 4.6 Thinking1512.8
  3. 03Claude Opus 4.7 Thinking1503.8
  4. 04Claude Opus 4.61498.9
  5. 05Claude Opus 4.8 Thinking1494.2
  6. 06Claude Opus 4.71492.9
  7. 07Claude Opus 4.5 Thinking 32K (2025-11-01)1483.9
  8. 08GPT-5.6 Sol X-High1483.5
  9. 09Gemini 3.1 Pro Preview1481.2
  10. 10Muse Spark 1.11478.1
Open the full instruction following arena board →
Hard prompts arenaTop 10 of 324 resolved rows
  1. 01Claude Fable 51533.4
  2. 02Claude Opus 4.6 Thinking1532.2
  3. 03Claude Opus 4.61526.7
  4. 04Claude Opus 4.7 Thinking1525.9
  5. 05Claude Opus 4.71519.1
  6. 06Claude Opus 4.8 Thinking1514.3
  7. 07Muse Spark 1.11513.6
  8. 08Gemini 3.1 Pro Preview1506.4
  9. 09GPT-5.6 Sol X-High1506.1
  10. 10Claude Opus 4.81505.8
Open the full hard prompts arena board →
Multi-turn arenaTop 10 of 322 resolved rows
  1. 01Claude Opus 4.71518.8
  2. 02Claude Opus 4.6 Thinking1517.9
  3. 03Claude Opus 4.7 Thinking1517.5
  4. 04Claude Fable 51514.9
  5. 05Claude Opus 4.61510.6
  6. 06Claude Opus 4.8 Thinking1506.3
  7. 07Muse Spark 1.11505.2
  8. 08Claude Opus 4.81499.4
  9. 09Gemini 3 Pro1495.3
  10. 10GPT-5.2 Chat Latest (2026-02-10)1494.4
Open the full multi-turn arena board →
Long queries arenaTop 10 of 319 resolved rows
  1. 01Claude Opus 4.6 Thinking1524.0
  2. 02Claude Fable 51524.0
  3. 03Claude Opus 4.7 Thinking1517.0
  4. 04Claude Opus 4.61517.0
  5. 05Claude Opus 4.71508.8
  6. 06Claude Opus 4.8 Thinking1506.2
  7. 07Claude Opus 4.81500.4
  8. 08Qwen3.7 Max Preview1497.8
  9. 09Gemini 3.1 Pro Preview1497.7
  10. 10Claude Opus 4.5 Thinking 32K (2025-11-01)1494.8
Open the full long queries arena board →
Math arenaTop 10 of 316 resolved rows
  1. 01Claude Fable 51550.2
  2. 02Claude Opus 4.6 Thinking1518.8
  3. 03Gemini 3.5 Flash High1517.6
  4. 04Claude Opus 4.7 Thinking1504.5
  5. 05Grok 4.51503.5
  6. 06Claude Opus 4.61502.1
  7. 07GPT-5.51496.7
  8. 08Claude Opus 4.8 Thinking1496.3
  9. 09GPT-5.4 High1494.9
  10. 10Claude Opus 4.71493.4
Open the full math arena board →
Vision arenaTop 10 of 108 resolved rows
  1. 01Claude Fable 51335.1
  2. 02Claude Opus 4.7 Thinking1318.4
  3. 03Claude Opus 4.61316.5
  4. 04Claude Opus 4.71315.5
  5. 05Claude Opus 4.6 Thinking1315.0
  6. 06Gemini 3.5 Flash Medium1308.8
  7. 07Muse Spark1307.3
  8. 08Gemini 3 Pro1304.9
  9. 09Gemini 3.5 Flash High1300.5
  10. 10GPT-5.4 High1298.8
Open the full vision arena board →
Web development arenaTop 10 of 56 resolved rows
  1. 01Claude Opus 4.7 Thinking1567.9
  2. 02Claude Opus 4.71561.9
  3. 03Claude Opus 4.6 Thinking1548.8
  4. 04Claude Opus 4.61544.4
  5. 05GLM-5.11531.7
  6. 06Claude Sonnet 4.61526.2
  7. 07Kimi K2.61524.6
  8. 08Muse Spark1509.4
  9. 09Claude Opus 4.5 Thinking 32K (2025-11-01)1490.6
  10. 10Mimo V2.5 Pro1476.4
Open the full web development arena board →
Search arenaTop 7 of 7 resolved rows
  1. 01Claude Opus 4.71234.9
  2. 02Claude Fable 51234.3
  3. 03Ernie 5.11227.0
  4. 04Claude Opus 4.81207.7
  5. 05Grok 4.20 Multi Agent Beta 03091206.9
  6. 06Grok 4.20 Beta 11189.5
  7. 07Grok 4.31163.1
Open the full search arena board →
Document arenaTop 10 of 32 resolved rows
  1. 01Claude Opus 4.6 Thinking1508.1
  2. 02Claude Fable 51507.1
  3. 03Claude Opus 4.61506.7
  4. 04Claude Opus 4.7 Thinking1504.5
  5. 05Claude Opus 4.71501.2
  6. 06GPT-5.5 High1488.2
  7. 07Claude Sonnet 4.61485.9
  8. 08GPT-5.51480.6
  9. 09Claude Opus 4.8 Thinking1474.7
  10. 10GPT 5.41472.2
Open the full document arena board →

Aider polyglot · code editing

Code editing systems, ranked by their own leaderboard.

Aider ranks full model-and-editor system configurations on code editing, not bare models. Pass rate 2 is Aider's own metric, and these rows are never compared against LMArena ratings.

  1. 01gpt-5 (high)88.0%
  2. 02gpt-5 (medium)86.7%
  3. 03o3-pro (high)84.9%
  4. 04gemini-2.5-pro-preview-06-05 (32k think)83.1%
  5. 05o3 (high)81.3%
  6. 06gpt-5 (low)81.3%
  7. 07grok-4 (high)79.6%
  8. 08gemini-2.5-pro-preview-06-05 (default think)79.1%
  9. 09o3 (high) + gpt-4.178.2%
  10. 10Gemini 2.5 Pro Preview 05-0676.9%
  11. 11o376.9%
  12. 12DeepSeek-V3.2-Exp (Reasoner)74.2%

Agent Arena · agentic tasks

The agent board, scored by its own arena.

LMArena's Agent Arena ranks models on real agentic sessions and published this board July 20, 2026.IPS is this arena's own score with its own scale. Rows labeled with a reasoning configuration are system results for that configuration, and a score at or below zero draws no bar.

  1. 01Claude Fable 5 (High)0.132
  2. 02Claude Opus 4.8 (Thinking)0.100
  3. 03GPT 5.6 Sol (xHigh)0.099
  4. 04Kimi K30.096
  5. 05Claude Sonnet 5 (High)0.091
  6. 06GPT 5.5 (xHigh)0.088
  7. 07Claude Opus 4.7 (Thinking)0.083
  8. 08Claude Opus 4.70.081
  9. 09GPT 5.5 (High)0.077
  10. 10Claude Opus 4.60.067
  11. 11GLM 5.2 (Max)0.065
  12. 12GPT 5.50.064

See all 37 agent rows →

EQ-Bench 3 · emotional intelligence

One source's judge-scored emotional intelligence board.

EQ-Bench 3 publishes its own Elo from judged role-play scenarios. The scale belongs to this source alone and shares nothing with LMArena numbers.

  1. 01Claude Opus 4.71461
  2. 02Claude Sonnet 4.61384
  3. 03Claude Opus 4.61383
  4. 04HiveLabsAI/hivemind 32b Preview1351
  5. 05Deepseek/deepseek V4 Pro1330
  6. 06Openai/gpt 5.41328
  7. 07Openai/gpt 5.51328
  8. 08Moonshotai/kimi K2.61319
  9. 09GPT 5.1 2025 11 131314
  10. 10GPT 5.21314
  11. 11Moonshotai/Kimi K2 Instruct1311
  12. 12Z Ai/glm 5.11310

Epoch AI Benchmarking Hub

Ten benchmarks, reported one at a time.

Every card is one benchmark's own top scores from Epoch AI's published hub data. Rows labeled with a scaffold, such as a coding agent or edit format, are system results, not bare-model results. Fraction-unit scores are shown as percentages.

Aider Polyglot

Aider Polyglot leaderboard · 69 resolved scores

  1. 01GPT 5 (2025-08-07, high) · diff88.0%
  2. 02GPT 5 (2025-08-07, medium) · diff86.7%
  3. 03O3 Pro (2025-06-10, high) · diff84.9%
  4. 04Gemini 2.5 Pro Preview 06 05 (32K) · diff-fenced83.1%
  5. 05GPT 5 (2025-08-07, low) · diff81.3%
  6. 06O3 (2025-04-16, high) · diff81.3%
  7. 07Grok 4 0709 · diff79.6%
  8. 08Grok 4 0709 (high) · diff79.6%

Chess Puzzles

Epoch AI Benchmarking Hub · 54 resolved scores

  1. 01GPT 5.5 Pro Pre Release (xhigh)64.0%
  2. 02GPT 5.6 Sol (promax)64.0%
  3. 03GPT 5.4 Pro (2026-03-05, xhigh)58.6%
  4. 04Gemini 3.1 Pro Preview55.0%
  5. 05GPT 5.6 Sol (max)55.0%
  6. 06GPT 5.5 Pre Release (xhigh)54.0%
  7. 07GPT 5.6 Terra (max)54.0%
  8. 08GPT 5.2 (2025-12-11, xhigh)49.0%

FrontierMath

Epoch AI Benchmarking Hub · 98 resolved scores

  1. 01GPT 5.5 Pro Pre Release (high)52.4%
  2. 02GPT 5.5 Pre Release (xhigh)51.7%
  3. 03GPT 5.5 Pro Pre Release (xhigh)51.0%
  4. 04GPT 5.4 Pro (2026-03-05, xhigh)50.0%
  5. 05GPT 5.4 (2026-03-05, xhigh)47.6%
  6. 06Claude Opus 4.8 (max)47.2%
  7. 07Claude Opus 4.7 (xhigh)43.8%
  8. 08Claude Opus 4.6 (max)40.7%

FrontierMath Tier 4

Epoch AI Benchmarking Hub · 69 resolved scores

  1. 01Gdm AI Co Mathematician47.9%
  2. 02GPT 5.5 Pro Pre Release (high)39.6%
  3. 03GPT 5.5 Pro Pre Release (xhigh)39.6%
  4. 04GPT 5.4 Pro 2026 03 05 Web App37.5%
  5. 05GPT 5.5 Pre Release (xhigh)35.4%
  6. 06GPT 5.2 Pro 2025 12 11 Webapp31.3%
  7. 07Claude Opus 4.8 (max)31.3%
  8. 08GPT 5.4 (2026-03-05, xhigh)27.1%

GPQA Diamond

Epoch AI Benchmarking Hub · 176 resolved scores

  1. 01GPT 5.4 Pro (2026-03-05, xhigh)94.6%
  2. 02Gemini 3.1 Pro Preview94.1%
  3. 03GPT 5.5 Pre Release (xhigh)94.0%
  4. 04GPT 5.5 Pro Pre Release (xhigh)93.9%
  5. 05GPT 5.6 Sol (max)93.5%
  6. 06Grok 4.5 (high)93.4%
  7. 07GPT 5.6 Terra (max)93.3%
  8. 08GPT 5.4 (2026-03-05, xhigh)93.3%

MATH Level 5

Epoch AI Benchmarking Hub · 106 resolved scores

  1. 01GPT 5 (2025-08-07, high)98.1%
  2. 02GPT 5 (2025-08-07, medium)97.9%
  3. 03GPT 5 Mini (2025-08-07, high)97.8%
  4. 04O4 Mini (2025-04-16, high)97.8%
  5. 05O3 (2025-04-16, high)97.8%
  6. 06Claude Sonnet 4 5 20250929 (32K)97.7%
  7. 07Qwen3 Max (2025-09-23)97.1%
  8. 08GPT 5 Mini (2025-08-07, medium)96.8%

OTIS Mock AIME 2024-2025

Epoch AI Benchmarking Hub · 150 resolved scores

  1. 01GPT 5.5 Pre Release (xhigh)100.0%
  2. 02GPT 5.5 Pro Pre Release (xhigh)100.0%
  3. 03GPT 5.6 Sol (max)100.0%
  4. 04Claude Fable 5 (max)99.7%
  5. 05GPT 5.6 Terra (max)99.7%
  6. 06Claude Opus 4.8 (max)98.3%
  7. 07GPT 5.6 Luna (max)98.3%
  8. 08Claude Opus 4.7 (xhigh)97.8%

SimpleQA Verified

Epoch AI Benchmarking Hub · 62 resolved scores

  1. 01Gemini 3.1 Pro Preview77.3%
  2. 02Gemini 3 Pro Preview72.9%
  3. 03GPT 5.6 Sol (max)71.6%
  4. 04Claude Fable 5 (xhigh)68.3%
  5. 05Qwen3 Max (2025-09-23)67.5%
  6. 06Gemini 3 Flash Preview67.4%
  7. 07Muse Spark66.3%
  8. 08GPT 5.5 Pro Pre Release (xhigh)64.5%

SWE-bench Verified

Epoch AI Benchmarking Hub · 33 resolved scores

  1. 01Claude Opus 4.7 (max)83.5%
  2. 02GPT 5.5 Pre Release (xhigh)80.6%
  3. 03Claude Opus 4.678.7%
  4. 04DeepSeek V4 Pro (max)77.6%
  5. 05Qwen3.7 Max77.3%
  6. 06GPT 5.4 (2026-03-05, high)76.9%
  7. 07Qwen3.6 Max Preview76.7%
  8. 08Claude Opus 4.5 (2025-11-01)76.7%

Terminal-Bench

Terminal-Bench leaderboard · 178 resolved scores

  1. 01Claude Opus 4.7 (unknown) · vix90.2%
  2. 02GPT-5.5 (unknown) · NexAU-AHE84.7%
  3. 03GPT-5.5 (unknown) · Capy83.1%
  4. 04GPT-5.5 (unknown) · Codex82.0%
  5. 05GPT-5.5 (unknown) · Codex CLI82.0%
  6. 06GPT 5.4 (2026-03-05, unknown) · ForgeCode81.8%
  7. 07Claude Opus 4.7 (unknown) · WOZCODE80.2%
  8. 08Gemini 3.1 Pro Preview · TongAgents80.2%

OpenRouter platform view

What the platform actually routes, lists, and charges.

OpenRouter's own usage rankings, model listings, and routed prices. Usage is the platform's weekly routed token volume,retrieved July 22, 2026: a live popularity signal that shows which models developers route work through right now, not a quality score. The platform's embedded third-party benchmark data is excluded at collection with a receipt and never appears on this site.

Most-used models this week

OpenRouter's routed token volume, week view. Volume is not a quality signal.

  1. 01Tencent: Hy310.9T tokens
  2. 02Xiaomi: MiMo-V2.59.4T tokens
  3. 03DeepSeek-V4-Flash5.4T tokens
  4. 04Z.ai: GLM 5.23.6T tokens
  5. 05MiniMax: MiniMax M33.2T tokens
  6. 06DeepSeek-V4-Pro2.8T tokens
  7. 07NVIDIA: Nemotron 3 Ultra2.7T tokens
  8. 08Claude Opus 4.71.9T tokens
  9. 09Claude Opus 4.81.9T tokens
  10. 10Anthropic: Claude Sonnet 51.1T tokens
  11. 11Google: Gemini 3 Flash Preview962.5B tokens
  12. 12StepFun: Step 3.7 Flash960.6B tokens
  13. 13anthropic/claude-4.6-sonnet876.8B tokens
  14. 14MoonshotAI: Kimi K3740.6B tokens
  15. 15GPT-5.5599.0B tokens

See the full usage board →

Newest listings

By OpenRouter's listing date. A listing date is not a release date.

  1. Poolside: Laguna S 2.12026-07-21

    1,048,576 token context$0.10 in / $0.20 out per 1M

  2. Google: Gemini 3.6 Flash2026-07-21

    1,048,576 token context$1.50 in / $7.50 out per 1M

  3. Google: Gemini 3.5 Flash-Lite2026-07-21

    1,048,576 token context$0.30 in / $2.50 out per 1M

  4. Meituan: LongCat 2.02026-07-20

    1,048,756 token context$0.30 in / $1.20 out per 1M

  5. Thinking Machines: Inkling2026-07-17

    524,288 token context$1.00 in / $4.05 out per 1M

  6. Auto Router (Beta)2026-07-17

    2,000,000 token contextVariable or unlisted pricing

  7. MoonshotAI: Kimi K32026-07-16

    1,048,576 token context$3.00 in / $15.00 out per 1M

  8. Meta: Muse Spark 1.12026-07-16

    1,048,576 token context$1.25 in / $4.25 out per 1M

  9. Kwaipilot: KAT-Coder-Air V2.52026-07-10

    256,000 token context$0.15 in / $0.60 out per 1M

  10. Kwaipilot: KAT-Coder-Pro V2.52026-07-10

    256,000 token context$0.74 in / $2.96 out per 1M

Frontier watchlist, priced

Input price per million tokens, cheapest first. Price says nothing about quality.

  1. 01DeepSeek V4 Pro$0.43 in / $0.87 out
  2. 02Gemini 3.1 Pro Preview$2.00 in / $12.00 out
  3. 03Grok 4.5$2.00 in / $6.00 out
  4. 04Claude Fable 5$10.00 in / $50.00 out

From the archive

How Meta moved through the historical AI race.

Open Meta’s race file
827.01395.91288.0
Legacy LMArena archive, May 8, 2023 to August 29, 2025

827 → 1396 → 1288

Meta’s reviewed historical path

This history includes Meta’s own models and excludes third-party models built from Llama.

Research and commercial inquiries

Work with the team behind the research.

Research questions, source submissions, corrections, or partnerships. Commercial relationships never influence a score, ranking, or editorial conclusion.

Direct answers

How Model Gauntlet works.

What is Model Gauntlet?

Model Gauntlet publishes sourced AI model standings and a separately labeled historical race archive. Current and legacy metrics are never blended.

How current are the standings?

The current provider table uses LMArena's published text style-control overall snapshot, dated July 16, 2026. It includes confidence intervals and battle counts.

Does the current table continue the historical chart?

No. The current LMArena rating and the legacy May 2023 to August 2025 archive use incompatible methodology and inclusion rules. No cross-window change is calculated.

Race alerts

Get the important moves in the AI race.

Join the waitlist for sourced updates on leaderboard changes, major model releases, and new research. The briefing launches after a final editorial review; early subscribers get the first edition.