← Model explorer

Sourced model profile · Current through July 21, 2026

Grok 4

A closed-weight xAI API model documented across language, multimodal, and vision domains, including search, OCR, and image-captioning tasks.

Developed byxAI
Released
July 9, 2025

What the evidence says

These facts come from the reviewed Epoch AI catalog artifact and its linked primary record. They describe the model. They do not rank it.

Developer [1]
xAI
Released [1]
July 9, 2025
Domains [1]
Language · Multimodal · Vision
Tasks [1]
Language modeling/generation · Question answering · Search · Visual question answering · Character recognition (OCR) · Image captioning · Quantitative reasoning
Access [1]
API access
Weights [1]
Closed

Where it may fit

These are evidence-bounded screening suggestions, not performance claims or purchase recommendations.

Mixed-media question answering

Consider it when a hosted workflow spans visual question answering, image captioning, and character recognition alongside language tasks.

Evidence fields: domains, tasks [1]

Search-connected API workflows

Consider it when search is an explicit documented task and API access fits the deployment.

Evidence fields: tasks, access [1]

Limitations

The record does not report retrieval quality, OCR accuracy, price, context window, latency, or safety evaluations, and no exact-match Arena rating exists for it. Listed tasks indicate intended scope, not measured performance.

Current standing

According to the LMArena snapshot published July 16, 2026, no exact match for this catalog record appears in the style-controlled text category, so this page reports no current rating for Grok 4.

Across the boards

Each card below is one named authority's own current view of Grok 4. Model Gauntlet mirrors these sources with attribution and does not combine them into a single score or rank Grok 4 across them.

Aider polyglot

Code editing systems

Pass rate
79.6%
Source rank
#7
System configuration
grok-4 (high)
Edit format
diff

Aider ranks the full model-and-editor-system configuration on polyglot code editing, not the bare model. This is the model portion of that system, never a bare-model Model Gauntlet score.

Source last updated May 22, 2026. Aider polyglot leaderboard, Aider-AI/aider, Apache-2.0.

Aider leaderboard source ↗

EQ-Bench 3

Emotional intelligence

Judge Elo
1190.3
Source rank
#42
Source label
grok-4

EQ-Bench 3 is a single-source subjective judge-based pairwise Elo for emotional-intelligence roleplays, appropriate only for that lane.

Source last updated May 10, 2026. EQ-Bench 3 canonical Elo leaderboard, EQ-bench/eqbench3, MIT.

EQ-Bench 3 source ↗