How Model Gauntlet works

How we report rankings without inventing our own score.

Model Gauntlet aggregates established research sources and explains what each one can show. The current leaderboard reports a single LMArena rating dated July 16, 2026, not a consensus ranking across research organizations. A separate historical archive contains 11 selected snapshots ending August 29, 2025. Because the methods changed, we do not combine the two views into one trend line.

Publication rules

Every ranking names its source, measure, and date.

We report what named authorities publish. Before publication, we check that names, values, dates, units, labels, and links match the source and that different methods are not accidentally blended. We do not rerun the authority's underlying research.

  • Current rating: one saved LMArena source file. It stays labeled with its published date until a newer snapshot passes human review.
  • Historical archive: eleven reviewed source files through August 29, 2025. It is not described as current.
  • New data: everything we pull from Epoch and LMArena goes through human review before it appears on the site. An automated download alone never updates a page.
  • Stale or unavailable inputs: the last reviewed publication may remain visible with its original date. It is never relabeled as freshly fetched.
  • Material changes: affected pages receive a modified date or correction linked to the public log. Silent factual rewrites are not permitted.

Formula, source, inclusion, or compatibility changes require a new methodology version and a dated change note.

Publication boundaries

What affects the ranking, and what does not.

Each source has a defined role. Only the current LMArena rating affects the leaderboard shown here.

ArchiveValidatePublish
01

Capability

The current table selects the highest Arena Score for each organization named in LMArena's published data, in one published, style-controlled LMArena rating for text models.

Current single-source rating
02

Usage data

No usage or adoption measure currently affects the leaderboard. A potential OpenRouter source remains excluded until each publishable field is confirmed as permitted.

Not published
03

Release velocity

Matched Epoch notable-model records provide dated release context. They do not affect any score.

Context only

Leaderboard selection

One leading model per organization.

For the current snapshot, the organization field comes from LMArena’s published dataset. We then select that organization’s highest Arena Score in the overall category. For historical records, we use a reviewed ownership list. If ownership is unclear or a model is a third-party derivative, we exclude it rather than guess.

frontier(organization, snapshot) = max(Arena Score in overall category)

Composite status

We do not publish a composite score.

The current leaderboard reports one LMArena measure. It is not a consensus ranking, and we do not combine capability, adoption, release activity, cost, or other factors into a proprietary score.

Any future composite needs a versioned formula, a source for every field, licensed inputs, fallbacks, a change log, and human approval before it can appear on an indexed page.

Update cadence

Why the historical chart stops in August 2025.

The historical chart contains 11 selected snapshots from May 8, 2023 through August 29, 2025. It crosses known LMArena method changes: Bradley-Terry replaced Elo on January 9, 2024; style control became the default on May 16, 2025; and frequency reweighting arrived July 23, 2025. The chart therefore describes published sampled ratings, not a normalized like-for-like trajectory. We do not append the July 2026 leaderboard because its scores are not sufficiently comparable with the historical series.

The evidence-gap explainercovers why no chart data exists between August 29, 2025 and July 16, 2026, with the source receipts for both eras.

Visibility measurement

RankPrompt measures discovery, not model quality.

Model Gauntlet runs 25 locked, unbranded US-English queries across AI race, current standings, model selection, comparison, and evidence clusters. The fixed engine set is ChatGPT, Claude, Perplexity, and Gemini.

RankPrompt is affiliated with the Anderson Collaborative ecosystem. These measurements guide publishing and outreach. They never supply or validate model scores, rankings, or research findings.

First complete baseline · Captured July 15, 2026

A transparent zero starting point.

First live US-English baseline completed across the locked 25-query pack. Zero Model Gauntlet mentions or owned-domain citations were observed; this is the starting measurement, not a quality score or model-ranking input.

Completeness
100/100
Mention rate
0.0%
Owned citation rate
0.0%
Engines
4 fixed
Visibility by query cluster. Each cluster contains five queries across four engines.
Query clusterCellsMention rateOwned citation rate
race200.0%0.0%
standings200.0%0.0%
selection200.0%0.0%
comparison200.0%0.0%
evidence200.0%0.0%

This public summary excludes raw responses, response text, competing domains, unrelated account data, and query-level citation records.

Sources and licensing

Each source has one defined role.

  • LMArena leaderboard datasetfor the current table, pinned to revision afed939e10281b660a4369206ca505b2bf5e0208and licensed CC BY 4.0.
  • Historical LMArena Space source filesfor the 11 selected legacy snapshots. Their Apache-2.0 repository records are kept separate from the current dataset.
  • Epoch AI notable modelsfor dated release context only.
  • The general OpenRouter model endpoint is not fetched or published. It remains excluded until a reviewed field list proves that every retained value is independently permitted and attributable.
  • Artificial Analysis data is never ingested or republished.

Method questions

Direct answers, no implied guarantees.

How often does the site update?

The current LMArena rating was published July 16, 2026. The separate historical view contains 11 selected snapshots through August 29, 2025. Neither is described as live.

What are the composite weights?

There is no published composite. The current table reports one LMArena rating source, not a consensus ranking. Epoch release records add context but do not change the table.

Does Model Gauntlet use Artificial Analysis data?

No. Model Gauntlet does not ingest or republish Artificial Analysis data. Its public presentation may inform interface research only.