The Math Ranking evaluates the mathematical rigor and quantitative reasoning capabilities of large language models. Rather than focusing on simple pattern recognition, these benchmarks test performance on formal proofs, competition-level Olympiad problems, multi-step algebra, and discrete mathematics, where accuracy and logical consistency are critical.
By aggregating standardized scores from various sources this leaderboard highlights the models that perform best in mathematical problem-solving, multi-step reasoning, and numerical accuracy.
Looking for a balanced model across all domains? Mathematical performance contributes 20% to the composite score. Visit the Overall LLM Ranking to see how top math models compare in coding, agentic capabilities, and broader logical reasoning.
| Rank | Model Name | Relative Quality | Score |
|---|---|---|---|
| 1 | Claude Opus 5.5 |
|
100.00 |
| 2 | Gemini 3.8 Flash |
|
98.01 |
| 3 | Muse Spark 1.3 |
|
97.69 |
| 4 | Claude Fable 5.1 |
|
97.45 |
| 5 | Gemini 3.7 Flash |
|
97.08 |
| 6 | Claude Opus 5 |
|
96.93 |
| 7 | Claude Fable 5 |
|
96.23 |
| 8 | MiMo-V2.6-Pro |
|
94.30 |
| 9 | Gemini 3.6 Flash |
|
94.11 |
| 10 | Muse Spark 1.2 |
|
93.92 |
| 11 | Claude Opus 4.6 |
|
93.65 |
| 12 | Gemini 3.1 Pro |
|
93.62 |
| 13 | Claude Opus 4.7 |
|
92.88 |
| 14 | Claude Opus 4.8 |
|
92.79 |
| 15 | GLM-5.3 |
|
92.73 |
| 16 | GLM-5.2 |
|
92.68 |
| 17 | Qwen3.7 |
|
92.62 |
| 18 | GPT-5.5 |
|
92.61 |
| 19 | Qwen3.6 |
|
92.56 |
| 20 | Grok 4.5 |
|
92.35 |
| 21 | DeepSeek-V4-Pro |
|
92.19 |
| 22 | GPT-6 Sol |
|
92.06 |
| 23 | Kimi K3 |
|
92.01 |
| 24 | Muse Spark 1.1 |
|
91.68 |
| 25 | MiMo-V2.6-Flash |
|
91.46 |
| 26 | Claude Sonnet 5 |
|
90.92 |
| 27 | DeepSeek-V4.1-Flash |
|
90.87 |
| 28 | Grok-4.20 |
|
90.82 |
| 29 | GPT-6 Astra |
|
90.65 |
| 30 | GPT-5.4 |
|
90.53 |
| 31 | GPT-5.6 Sol |
|
90.52 |
| 32 | Hy4 |
|
90.52 |
| 33 | Grok 4.6 |
|
90.02 |
| 34 | Seed 2.1 Pro |
|
89.73 |
| 35 | Qwen3.7-Plus |
|
89.72 |
| 36 | Muse Spark |
|
89.49 |
| 37 | GPT-5.2 |
|
89.14 |
| 38 | Grok 4.7 |
|
88.97 |
| 39 | Hy3 |
|
88.78 |
| 40 | GLM-5.3-Flash |
|
88.70 |
| 41 | Qwen3.8 |
|
88.69 |
| 42 | MiMo-V2-Pro |
|
88.61 |
| 43 | Gemini 3 Pro |
|
88.46 |
| 44 | Gemini 3 Flash |
|
88.22 |
| 45 | Kimi K2.6 |
|
88.20 |
| 46 | Muse Glimmer |
|
88.13 |
| 47 | DeepSeek-V4-Flash |
|
87.97 |
| 48 | Qwen3.6 Plus |
|
87.94 |
| 49 | Kimi K2.5 |
|
87.67 |
| 50 | GPT-5.6 Terra |
|
87.47 |
| 51 | GPT-6 Luna |
|
87.44 |
| 52 | GLM-5 |
|
87.34 |
| 53 | Grok-4.1 |
|
87.18 |
| 54 | MiMo-V2.5 |
|
86.55 |
| 55 | GLM-5V-Turbo |
|
86.39 |
| 56 | Qwen3.5-397B-A17B |
|
86.17 |
| 57 | Claude Sonnet 4.6 |
|
86.16 |
| 58 | Claude Opus 4.5 |
|
86.10 |
| 59 | Gemini 3.5 Flash |
|
86.09 |
| 60 | MiMo-V2-Omni |
|
86.08 |
| 61 | GLM-5.1 |
|
85.22 |
| 62 | Grok 4.3 |
|
84.64 |
| 63 | Kimi K2 |
|
84.35 |
| 64 | Inkling |
|
84.24 |
| 65 | MiniMax M2.7 |
|
84.02 |
| 66 | Grok 4.1 Fast |
|
83.86 |
| 67 | Qwen3.5-122B-A10B |
|
83.72 |
| 68 | Inkling-Small |
|
83.64 |
| 69 | GPT-5.6 Luna |
|
83.36 |
| 70 | Gemini 3.5 Flash-Lite |
|
83.21 |
| 71 | Qwen3.8-27B |
|
83.09 |
| 72 | Qwen3.5-27B |
|
82.69 |
| 73 | GPT-5.1 |
|
82.61 |
| 74 | GLM-4.7 |
|
82.61 |
| 75 | Seed 2.0 Pro |
|
82.53 |
| 76 | Gemma 4 31B |
|
81.83 |
| 77 | ChatGPT-4o |
|
81.17 |
| 78 | Qwen3 |
|
81.11 |
| 79 | Qwen3.8-Flash-Next |
|
81.06 |
| 80 | Grok Build 0.1 |
|
80.74 |
| 81 | GPT-5 |
|
80.72 |
| 82 | Qwen3.6-27B |
|
80.61 |
| 83 | Step-3.5-Flash |
|
80.55 |
| 84 | ERNIE 5.0 |
|
80.40 |
| 85 | DeepSeek-V3.2 |
|
80.24 |
| 86 | DeepSeek V3.1 Terminus |
|
80.06 |
| 87 | LongCat-Flash |
|
80.00 |
| 88 | Grok-4 |
|
79.92 |
| 89 | Mistral |
|
79.91 |
| 90 | DeepSeek-V3.2-Speciale |
|
79.78 |
| 91 | MiMo-V2.5-Pro |
|
79.46 |
| 92 | MiniMax M2.5 |
|
79.43 |
| 93 | Qwen3 235B A22B |
|
79.21 |
| 94 | Qwen3.5-35B-A3B |
|
78.82 |
| 95 | Gemma 4 26B-A4B |
|
78.51 |
| 96 | Trinity Large |
|
77.85 |
| 97 | Qwen3 VL 235B A22B |
|
77.24 |
| 98 | Claude Sonnet 4.5 |
|
77.24 |
| 99 | o3 |
|
77.08 |
| 100 | Solar Pro 4 |
|
76.78 |