Coding
According to the Epoch AI model catalog (retrieved July 21, 2026), 75 reviewed records explicitly list code generation as a task: MiniMax-M2.1, GLM-4.7, Claude Opus 4.5, Olmo 3, Grok 4.1 Fast, Grok 4.1, Kimi K2 Thinking, MiniMax-M2, Claude Haiku 4.5, Claude Sonnet 4.5, AgentFounder-30B, Qwen3-Max, Claude Opus 4.1, Gemini 2.5 Deep Think, Qwen3-235B-A22B (Jul 2025), Qwen3-235B-A22B-Thinking (Jul 2025), Qwen3-Coder-480B-A35B, Kimi K2, Grok 4 Heavy, Gemini 2.5 Pro (Jun 2025), DeepSeek-R1 (May 2025), Claude Opus 4, Claude Sonnet 4, Gemini 2.5 Pro (May 2025), Qwen3-235B-A22B, Llama 4 Behemoth (preview), Llama 4 Maverick, Llama 4 Scout, Gemini 2.5 Pro (Mar 2025), DeepSeek-V3 (Mar 2025), o1-pro, ERNIE-4.5-VL-424B-A47B (文心大模型4.5), Hunyuan-TurboS, QwQ-32B, GPT-4.5, Claude 3.7 Sonnet, Grok 3, o3-mini, Kimi k1.5, DeepSeek-R1, DeepSeek-V3, o3, Gemini 2.0 Pro, Llama 3.3 70B, o1, Amazon Nova Pro, Hunyuan-Large, NVLM-D 72B, NVLM-H 72B, NVLM-X 72B, Qwen2.5 Instruct (72B), o1-mini, o1-preview, DeepSeek-V2.5, Grok-2, Mistral Large 2, Llama 3.1-405B, GPT-4o mini, Claude 3.5 Sonnet, DeepSeek-Coder-V2 236B, GLM-4 (0520), Llama 3-70B, Claude 3 Opus, Claude 3 Sonnet, Qwen1.5-72B, FunSearch, Qwen-72B, Nemotron-3-8B, Yi-34B, ChatGLM3-6B, CODEFUSION (Python), Amazon Titan, LLaMA-65B, PaLM (540B), and AlphaCode. That catalog does not measure coding quality, so no coding winner can be named from this evidence.
A task label establishes relevance, not performance. Run representative repository tasks for correctness, tool use, latency, and cost before choosing.
Open the coding shortlist guide →