The Math Ranking evaluates the mathematical rigor and quantitative reasoning capabilities of large language models. Rather than focusing on simple pattern recognition, these benchmarks test performance on formal proofs, competition-level Olympiad problems, multi-step algebra, and discrete mathematics, where accuracy and logical consistency are critical.

By aggregating standardized scores from various sources this leaderboard highlights the models that perform best in mathematical problem-solving, multi-step reasoning, and numerical accuracy.

Looking for a balanced model across all domains? Mathematical performance contributes 20% to the composite score. Visit the Overall LLM Ranking to see how top math models compare in coding, agentic capabilities, and broader logical reasoning.

Last updated: September 26, 2026
Rank Model Name Relative Quality Score
1 Claude Opus 5.5
Quality of Math 100.00%

100.00
2 Gemini 3.8 Flash
Quality of Math 98.01%

98.01
3 Muse Spark 1.3
Quality of Math 97.69%

97.69
4 Claude Fable 5.1
Quality of Math 97.45%

97.45
5 Gemini 3.7 Flash
Quality of Math 97.08%

97.08
6 Claude Opus 5
Quality of Math 96.93%

96.93
7 Claude Fable 5
Quality of Math 96.23%

96.23
8 MiMo-V2.6-Pro
Quality of Math 94.30%

94.30
9 Gemini 3.6 Flash
Quality of Math 94.11%

94.11
10 Muse Spark 1.2
Quality of Math 93.92%

93.92
11 Claude Opus 4.6
Quality of Math 93.65%

93.65
12 Gemini 3.1 Pro
Quality of Math 93.62%

93.62
13 Claude Opus 4.7
Quality of Math 92.88%

92.88
14 Claude Opus 4.8
Quality of Math 92.79%

92.79
15 GLM-5.3
Quality of Math 92.73%

92.73
16 GLM-5.2
Quality of Math 92.68%

92.68
17 Qwen3.7
Quality of Math 92.62%

92.62
18 GPT-5.5
Quality of Math 92.61%

92.61
19 Qwen3.6
Quality of Math 92.56%

92.56
20 Grok 4.5
Quality of Math 92.35%

92.35
21 DeepSeek-V4-Pro
Quality of Math 92.19%

92.19
22 GPT-6 Sol
Quality of Math 92.06%

92.06
23 Kimi K3
Quality of Math 92.01%

92.01
24 Muse Spark 1.1
Quality of Math 91.68%

91.68
25 MiMo-V2.6-Flash
Quality of Math 91.46%

91.46
26 Claude Sonnet 5
Quality of Math 90.92%

90.92
27 DeepSeek-V4.1-Flash
Quality of Math 90.87%

90.87
28 Grok-4.20
Quality of Math 90.82%

90.82
29 GPT-6 Astra
Quality of Math 90.65%

90.65
30 GPT-5.4
Quality of Math 90.53%

90.53
31 GPT-5.6 Sol
Quality of Math 90.52%

90.52
32 Hy4
Quality of Math 90.52%

90.52
33 Grok 4.6
Quality of Math 90.02%

90.02
34 Seed 2.1 Pro
Quality of Math 89.73%

89.73
35 Qwen3.7-Plus
Quality of Math 89.72%

89.72
36 Muse Spark
Quality of Math 89.49%

89.49
37 GPT-5.2
Quality of Math 89.14%

89.14
38 Grok 4.7
Quality of Math 88.97%

88.97
39 Hy3
Quality of Math 88.78%

88.78
40 GLM-5.3-Flash
Quality of Math 88.70%

88.70
41 Qwen3.8
Quality of Math 88.69%

88.69
42 MiMo-V2-Pro
Quality of Math 88.61%

88.61
43 Gemini 3 Pro
Quality of Math 88.46%

88.46
44 Gemini 3 Flash
Quality of Math 88.22%

88.22
45 Kimi K2.6
Quality of Math 88.20%

88.20
46 Muse Glimmer
Quality of Math 88.13%

88.13
47 DeepSeek-V4-Flash
Quality of Math 87.97%

87.97
48 Qwen3.6 Plus
Quality of Math 87.94%

87.94
49 Kimi K2.5
Quality of Math 87.67%

87.67
50 GPT-5.6 Terra
Quality of Math 87.47%

87.47
51 GPT-6 Luna
Quality of Math 87.44%

87.44
52 GLM-5
Quality of Math 87.34%

87.34
53 Grok-4.1
Quality of Math 87.18%

87.18
54 MiMo-V2.5
Quality of Math 86.55%

86.55
55 GLM-5V-Turbo
Quality of Math 86.39%

86.39
56 Qwen3.5-397B-A17B
Quality of Math 86.17%

86.17
57 Claude Sonnet 4.6
Quality of Math 86.16%

86.16
58 Claude Opus 4.5
Quality of Math 86.10%

86.10
59 Gemini 3.5 Flash
Quality of Math 86.09%

86.09
60 MiMo-V2-Omni
Quality of Math 86.08%

86.08
61 GLM-5.1
Quality of Math 85.22%

85.22
62 Grok 4.3
Quality of Math 84.64%

84.64
63 Kimi K2
Quality of Math 84.35%

84.35
64 Inkling
Quality of Math 84.24%

84.24
65 MiniMax M2.7
Quality of Math 84.02%

84.02
66 Grok 4.1 Fast
Quality of Math 83.86%

83.86
67 Qwen3.5-122B-A10B
Quality of Math 83.72%

83.72
68 Inkling-Small
Quality of Math 83.64%

83.64
69 GPT-5.6 Luna
Quality of Math 83.36%

83.36
70 Gemini 3.5 Flash-Lite
Quality of Math 83.21%

83.21
71 Qwen3.8-27B
Quality of Math 83.09%

83.09
72 Qwen3.5-27B
Quality of Math 82.69%

82.69
73 GPT-5.1
Quality of Math 82.61%

82.61
74 GLM-4.7
Quality of Math 82.61%

82.61
75 Seed 2.0 Pro
Quality of Math 82.53%

82.53
76 Gemma 4 31B
Quality of Math 81.83%

81.83
77 ChatGPT-4o
Quality of Math 81.17%

81.17
78 Qwen3
Quality of Math 81.11%

81.11
79 Qwen3.8-Flash-Next
Quality of Math 81.06%

81.06
80 Grok Build 0.1
Quality of Math 80.74%

80.74
81 GPT-5
Quality of Math 80.72%

80.72
82 Qwen3.6-27B
Quality of Math 80.61%

80.61
83 Step-3.5-Flash
Quality of Math 80.55%

80.55
84 ERNIE 5.0
Quality of Math 80.40%

80.40
85 DeepSeek-V3.2
Quality of Math 80.24%

80.24
86 DeepSeek V3.1 Terminus
Quality of Math 80.06%

80.06
87 LongCat-Flash
Quality of Math 80.00%

80.00
88 Grok-4
Quality of Math 79.92%

79.92
89 Mistral
Quality of Math 79.91%

79.91
90 DeepSeek-V3.2-Speciale
Quality of Math 79.78%

79.78
91 MiMo-V2.5-Pro
Quality of Math 79.46%

79.46
92 MiniMax M2.5
Quality of Math 79.43%

79.43
93 Qwen3 235B A22B
Quality of Math 79.21%

79.21
94 Qwen3.5-35B-A3B
Quality of Math 78.82%

78.82
95 Gemma 4 26B-A4B
Quality of Math 78.51%

78.51
96 Trinity Large
Quality of Math 77.85%

77.85
97 Qwen3 VL 235B A22B
Quality of Math 77.24%

77.24
98 Claude Sonnet 4.5
Quality of Math 77.24%

77.24
99 o3
Quality of Math 77.08%

77.08
100 Solar Pro 4
Quality of Math 76.78%

76.78