Updated Sept 2026

Reasoning leaderboard

Composite R-score across six suites. Click any column to sort; use the filters to compare open and closed models. Lower hallucination rate is better.

Rank Model Organization Overall Math Code Logic Multi-step Halluc. ↓ 30d Δ
1 Claude Opus 4.5 Anthropic 89.2
91.4 90.1 88.7 86.9 4.1 ▲ 0.8
2 GPT-5.2 OpenAI 88.6
92.0 88.9 87.5 86.0 5.2 ▲ 1.2
3 Gemini 3 Pro Google 87.9
89.8 87.2 89.0 85.6 4.8 ▲ 0.5
4 Grok 4.1 xAI 85.1
86.3 84.0 86.8 83.3 5.9 ▲ 0.4
5 DeepSeek-R2Open DeepSeek 84.7
87.9 86.4 82.1 82.4 6.3 ▲ 2.1
6 GLM-5Open Z.ai 84.0
85.2 85.8 82.9 82.1 6.6 ▲ 1.8
7 Qwen3-MaxOpen Alibaba 83.4
84.7 83.1 83.0 82.8 7.0 ▲ 0.9
8 Kimi K2.5Open Moonshot AI 82.6
83.4 84.2 81.0 81.8 7.4 ▲ 1.5
9 Llama 4 MaverickOpen Meta 78.9
77.5 79.8 80.2 78.1 9.2 ▲ 0.3
10 Mistral Large 3 Mistral AI 77.2
76.1 78.4 78.9 76.6 9.8 ▲ 0.6

10 / 10 models · Top 10 shown of 27 tracked. Scores are illustrative while the evaluation pipeline is finalized — ±95% bootstrap confidence intervals of ±0.9–1.6 omitted for readability.