The Reasoning Leaderboard shows how AI models able to handle complex logical and analytical tasks. It measures deep logical deduction, counterfactual reasoning, scientific problem-solving, and consistency across long, multi-turn interactions. These benchmarks are especially relevant for tasks that require more than simple information retrieval or pattern recognition.
Our composite score aggregates several leading reasoning benchmarks. Together, they highlight the models that perform best on complex reasoning, scientific analysis, and multi-step problem-solving.
Looking for the full picture? Reasoning accounts for 30% of the overall evaluation. Visit the Overall LLM Ranking to see how top reasoning models perform across coding, mathematics, and autonomous agent tasks.
| Rank | Model Name | Relative Quality | Score |
|---|---|---|---|
| 1 | Claude Opus 5.5 |
|
99.82 |
| 2 | GPT-6 Astra |
|
97.58 |
| 3 | Claude Fable 5.1 |
|
95.97 |
| 4 | Claude Fable 5 |
|
95.24 |
| 5 | Claude Opus 5 |
|
95.14 |
| 6 | GPT-5.6 Sol |
|
93.82 |
| 7 | Kimi K3 |
|
92.94 |
| 8 | Muse Spark 1.3 |
|
92.81 |
| 9 | Claude Opus 4.8 |
|
92.12 |
| 10 | Gemini 3.8 Flash |
|
91.38 |
| 11 | Gemini 3.7 Flash |
|
90.72 |
| 12 | Muse Spark 1.1 |
|
90.59 |
| 13 | GPT-5.5 |
|
90.56 |
| 14 | Qwen3.8 |
|
90.45 |
| 15 | Hy4 |
|
90.38 |
| 16 | GPT-5.6 Terra |
|
89.68 |
| 17 | GLM-5.3 |
|
89.56 |
| 18 | GPT-6 Sol |
|
89.44 |
| 19 | Gemini 3.1 Pro |
|
89.04 |
| 20 | Claude Opus 4.7 |
|
88.94 |
| 21 | Grok 4.6 |
|
88.90 |
| 22 | Claude Opus 4.6 |
|
88.60 |
| 23 | DeepSeek-V4-Pro |
|
88.58 |
| 24 | Claude Sonnet 5 |
|
88.04 |
| 25 | GPT-5.4 |
|
87.79 |
| 26 | Grok 4.5 |
|
87.11 |
| 27 | Qwen3.7 |
|
86.97 |
| 28 | Gemini 3.5 Flash |
|
86.16 |
| 29 | DeepSeek-V4.1-Flash |
|
86.09 |
| 30 | Seed 2.1 Pro |
|
86.00 |
| 31 | Qwen3.8-Flash-Next |
|
85.92 |
| 32 | GLM-5.3-Flash |
|
85.86 |
| 33 | Muse Spark 1.2 |
|
85.52 |
| 34 | Gemini 3.6 Flash |
|
85.42 |
| 35 | MiMo-V2.6-Pro |
|
85.03 |
| 36 | MiMo-V2.6-Flash |
|
85.01 |
| 37 | GPT-5.6 Luna |
|
84.94 |
| 38 | GLM-5.2 |
|
84.63 |
| 39 | GPT-5.2 |
|
84.36 |
| 40 | DeepSeek-V4-Flash |
|
84.30 |
| 41 | Muse Spark |
|
83.38 |
| 42 | Kimi K2.6 |
|
83.30 |
| 43 | Gemini 3 Pro |
|
82.94 |
| 44 | GPT-6 Luna |
|
82.31 |
| 45 | Qwen3.8-27B |
|
82.27 |
| 46 | Qwen3.7-Plus |
|
81.87 |
| 47 | MiniMax M3 |
|
81.75 |
| 48 | Claude Sonnet 4.6 |
|
81.69 |
| 49 | Seed 2.0 Pro |
|
81.69 |
| 50 | Gemini 3 Flash |
|
80.71 |
| 51 | Qwen3.6 |
|
80.59 |
| 52 | Claude Opus 4.5 |
|
80.54 |
| 53 | Hy3 |
|
80.36 |
| 54 | Grok 4.7 |
|
79.60 |
| 55 | Grok Build 0.1 |
|
79.02 |
| 56 | GLM-5.1 |
|
78.30 |
| 57 | LongCat-Flash |
|
77.91 |
| 58 | Qwen3.6 Plus |
|
77.90 |
| 59 | Inkling |
|
77.90 |
| 60 | Kimi K2.5 |
|
77.45 |
| 61 | GPT-5.3 Codex |
|
77.17 |
| 62 | Qwen3.5-397B-A17B |
|
77.05 |
| 63 | MiMo-V2-Pro |
|
77.01 |
| 64 | Inkling-Small |
|
76.35 |
| 65 | Kimi K2.7 Code |
|
76.20 |
| 66 | GLM-5 |
|
75.85 |
| 67 | GPT-5.1 |
|
75.81 |
| 68 | GPT-5.2 Codex |
|
75.30 |
| 69 | Step 5 |
|
75.00 |
| 70 | GPT-5.4 mini |
|
74.55 |
| 71 | MiMo-V2.5 |
|
74.13 |
| 72 | GLM-4.7 |
|
74.01 |
| 73 | MiniMax M2.7 |
|
74.00 |
| 74 | Qwen3 |
|
73.76 |
| 75 | MiMo-V2.5-Pro |
|
73.73 |
| 76 | GPT-5 |
|
73.72 |
| 77 | Grok 4.3 |
|
73.20 |
| 78 | Gemma 4 31B |
|
73.19 |
| 79 | ChatGPT-4o |
|
72.77 |
| 80 | Muse Glimmer |
|
72.26 |
| 81 | GPT-5.4 nano |
|
72.20 |
| 82 | Qwen3.5-122B-A10B |
|
72.14 |
| 83 | GPT-5.5 Instant |
|
72.11 |
| 84 | Kimi K2 |
|
71.75 |
| 85 | Grok-4 |
|
71.72 |
| 86 | Qwen3.5-27B |
|
71.24 |
| 87 | DeepSeek-V3.2 |
|
71.12 |
| 88 | KAT-Coder-Pro V2 |
|
71.10 |
| 89 | Step-3.5-Flash |
|
70.87 |
| 90 | Grok-4.20 |
|
70.35 |
| 91 | MiniMax M2.5 |
|
70.28 |
| 92 | Claude Sonnet 4.5 |
|
70.27 |
| 93 | Grok 4.1 Fast |
|
70.03 |
| 94 | GPT-4.5 |
|
69.84 |
| 95 | Grok-4.1 |
|
69.46 |
| 96 | MiMo-V2-Omni |
|
69.04 |
| 97 | Gemini 2.5 Pro |
|
68.85 |
| 98 | Solar Pro 4 |
|
68.77 |
| 99 | MiniMax M2.1 |
|
68.16 |
| 100 | Qwen3.6-27B |
|
68.15 |