Capability
The current table selects the highest Arena Score for each organization named in LMArena's published data, in one published, style-controlled LMArena rating for text models.
Current single-source ratingHow Model Gauntlet works
Model Gauntlet aggregates established research sources and explains what each one can show. The current leaderboard reports a single LMArena rating dated July 16, 2026, not a consensus ranking across research organizations. A separate historical archive contains 11 selected snapshots ending August 29, 2025. Because the methods changed, we do not combine the two views into one trend line.
Publication rules
We report what named authorities publish. Before publication, we check that names, values, dates, units, labels, and links match the source and that different methods are not accidentally blended. We do not rerun the authority's underlying research.
Formula, source, inclusion, or compatibility changes require a new methodology version and a dated change note.
Publication boundaries
Each source has a defined role. Only the current LMArena rating affects the leaderboard shown here.
The current table selects the highest Arena Score for each organization named in LMArena's published data, in one published, style-controlled LMArena rating for text models.
Current single-source ratingNo usage or adoption measure currently affects the leaderboard. A potential OpenRouter source remains excluded until each publishable field is confirmed as permitted.
Not publishedMatched Epoch notable-model records provide dated release context. They do not affect any score.
Context onlyLeaderboard selection
For the current snapshot, the organization field comes from LMArena’s published dataset. We then select that organization’s highest Arena Score in the overall category. For historical records, we use a reviewed ownership list. If ownership is unclear or a model is a third-party derivative, we exclude it rather than guess.
frontier(organization, snapshot) = max(Arena Score in overall category)Composite status
The current leaderboard reports one LMArena measure. It is not a consensus ranking, and we do not combine capability, adoption, release activity, cost, or other factors into a proprietary score.
Any future composite needs a versioned formula, a source for every field, licensed inputs, fallbacks, a change log, and human approval before it can appear on an indexed page.
Update cadence
The historical chart contains 11 selected snapshots from May 8, 2023 through August 29, 2025. It crosses known LMArena method changes: Bradley-Terry replaced Elo on January 9, 2024; style control became the default on May 16, 2025; and frequency reweighting arrived July 23, 2025. The chart therefore describes published sampled ratings, not a normalized like-for-like trajectory. We do not append the July 2026 leaderboard because its scores are not sufficiently comparable with the historical series.
The evidence-gap explainercovers why no chart data exists between August 29, 2025 and July 16, 2026, with the source receipts for both eras.
Visibility measurement
Model Gauntlet runs 25 locked, unbranded US-English queries across AI race, current standings, model selection, comparison, and evidence clusters. The fixed engine set is ChatGPT, Claude, Perplexity, and Gemini.
RankPrompt is affiliated with the Anderson Collaborative ecosystem. These measurements guide publishing and outreach. They never supply or validate model scores, rankings, or research findings.
First complete baseline · Captured July 15, 2026
First live US-English baseline completed across the locked 25-query pack. Zero Model Gauntlet mentions or owned-domain citations were observed; this is the starting measurement, not a quality score or model-ranking input.
| Query cluster | Cells | Mention rate | Owned citation rate |
|---|---|---|---|
| race | 20 | 0.0% | 0.0% |
| standings | 20 | 0.0% | 0.0% |
| selection | 20 | 0.0% | 0.0% |
| comparison | 20 | 0.0% | 0.0% |
| evidence | 20 | 0.0% | 0.0% |
This public summary excludes raw responses, response text, competing domains, unrelated account data, and query-level citation records.
Sources and licensing
afed939e10281b660a4369206ca505b2bf5e0208and licensed CC BY 4.0.Method questions
The current LMArena rating was published July 16, 2026. The separate historical view contains 11 selected snapshots through August 29, 2025. Neither is described as live.
There is no published composite. The current table reports one LMArena rating source, not a consensus ranking. Epoch release records add context but do not change the table.
No. Model Gauntlet does not ingest or republish Artificial Analysis data. Its public presentation may inform interface research only.