The Coding Leaderboard evaluates LLMs on their software engineering capabilities. From generating production-ready code across major programming languages to designing algorithms, debugging complex functions, and translating scientific concepts into executable scripts, this category measures a model’s ability to handle a broad range of coding and software development tasks.
Scores in this category are compiled from leading coding benchmarks highlighting models that perform well on code generation, debugging, problem-solving, and technical accuracy.
Need more than code generation? Coding performance accounts for 30% of our composite rating. Explore the Overall LLM Ranking to see which top coding models also excel at autonomous agent workflows, advanced mathematics, and complex reasoning.
| Rank | Model Name | Relative Quality | Score |
|---|---|---|---|
| 1 | Claude Opus 5.5 |
|
99.88 |
| 2 | GPT-6 Astra |
|
93.89 |
| 3 | GPT-5.6 Sol |
|
93.37 |
| 4 | ERNIE 5.0 |
|
90.87 |
| 5 | Claude Fable 5.1 |
|
90.09 |
| 6 | Claude Fable 5 |
|
89.81 |
| 7 | GPT-5.6 Terra |
|
88.17 |
| 8 | Kimi K3 |
|
88.05 |
| 9 | Claude Opus 5 |
|
87.82 |
| 10 | Muse Spark 1.3 |
|
87.62 |
| 11 | GPT-6 Sol |
|
87.48 |
| 12 | GPT-5.5 |
|
87.44 |
| 13 | ChatGPT-4o |
|
87.43 |
| 14 | Qwen3 VL 235B A22B |
|
86.83 |
| 15 | MiMo-V2.6-Flash |
|
86.51 |
| 16 | GPT-5.6 Luna |
|
85.90 |
| 17 | MiMo-V2.6-Pro |
|
85.72 |
| 18 | GLM-5.3 |
|
85.49 |
| 19 | GPT-5.5 Instant |
|
85.49 |
| 20 | Mistral |
|
85.48 |
| 21 | DeepSeek-V4.1-Flash |
|
84.88 |
| 22 | Grok 4.7 |
|
83.57 |
| 23 | Qwen3.8 |
|
83.19 |
| 24 | Claude Opus 4.8 |
|
82.57 |
| 25 | Claude Opus 4.7 |
|
81.79 |
| 26 | Gemini 3.7 Flash |
|
81.78 |
| 27 | Grok 4.6 |
|
81.72 |
| 28 | GPT-6 Luna |
|
81.47 |
| 29 | Gemini 3.8 Flash |
|
80.84 |
| 30 | DeepSeek-V4-Pro |
|
80.55 |
| 31 | GLM-5.3-Flash |
|
80.32 |
| 32 | GLM-5.2 |
|
80.13 |
| 33 | INTELLECT-3 |
|
79.79 |
| 34 | GLM-4.6V |
|
79.79 |
| 35 | Claude Sonnet 5 |
|
79.68 |
| 36 | Muse Spark 1.1 |
|
79.68 |
| 37 | Hy4 |
|
79.32 |
| 38 | Mistral Small |
|
79.19 |
| 39 | Ling-flash-2.0 |
|
78.89 |
| 40 | Claude Opus 4.6 |
|
78.52 |
| 41 | GLM-4.5V |
|
77.69 |
| 42 | Qwen2.5 |
|
77.69 |
| 43 | Qwen3.8-Flash-Next |
|
77.35 |
| 44 | DeepSeek-V4-Flash |
|
77.23 |
| 45 | Muse Spark 1.2 |
|
77.05 |
| 46 | Grok 4.5 |
|
76.77 |
| 47 | Kimi K2.6 |
|
76.46 |
| 48 | Qwen3.7 |
|
76.33 |
| 49 | Qwen3.8-27B |
|
76.11 |
| 50 | Ring-flash-2.0 |
|
75.90 |
| 51 | Gemini 3.6 Flash |
|
75.45 |
| 52 | Gemini 3.5 Flash |
|
75.39 |
| 53 | Qwen3.7-Plus |
|
75.32 |
| 54 | GPT-5 nano |
|
75.00 |
| 55 | Gemini 3.1 Pro |
|
74.77 |
| 56 | Olmo 3.1 32B |
|
74.55 |
| 57 | Claude Opus 4.5 |
|
74.32 |
| 58 | Claude Sonnet 4.6 |
|
73.85 |
| 59 | Muse Spark |
|
73.39 |
| 60 | Grok Build 0.1 |
|
73.24 |
| 61 | Step 5 |
|
73.22 |
| 62 | Hy3 |
|
71.92 |
| 63 | MiMo-V2.5-Pro |
|
71.76 |
| 64 | Qwen3.6 Plus |
|
71.61 |
| 65 | Olmo 3 32B Think |
|
71.41 |
| 66 | Seed 2.0 Pro |
|
71.40 |
| 67 | GPT-5.4 |
|
70.91 |
| 68 | MiniMax M3 |
|
70.88 |
| 69 | Mistral Large |
|
70.51 |
| 70 | GPT-5.3 Codex |
|
70.03 |
| 71 | GLM-5.1 |
|
69.73 |
| 72 | Qwen3.6 |
|
69.05 |
| 73 | DeepSeek V3.1 Terminus |
|
69.02 |
| 74 | Olmo 3.1 32B Think |
|
68.56 |
| 75 | MiMo-V2-Omni |
|
68.42 |
| 76 | Qwen3.6-27B |
|
68.29 |
| 77 | Kimi K2.7 Code |
|
68.18 |
| 78 | GPT-5.2 |
|
67.86 |
| 79 | Kimi K2 |
|
67.78 |
| 80 | GPT-5.4 nano |
|
67.68 |
| 81 | GPT-5 mini |
|
67.68 |
| 82 | LongCat-Flash |
|
67.52 |
| 83 | Seed 2.1 Pro |
|
67.23 |
| 84 | Inkling |
|
67.17 |
| 85 | Llama 3.1 Nemotron 70B |
|
66.62 |
| 86 | Step-3.5-Flash |
|
66.28 |
| 87 | GPT-5.4 mini |
|
66.13 |
| 88 | Claude Opus 4 |
|
65.97 |
| 89 | Inkling-Small |
|
65.38 |
| 90 | MiMo-V2.5 |
|
65.35 |
| 91 | Grok 4.3 |
|
65.15 |
| 92 | MiniMax M2.7 |
|
64.47 |
| 93 | Granite 4.2 30B |
|
64.38 |
| 94 | GLM-5 |
|
64.30 |
| 95 | MiMo-V2-Pro |
|
64.24 |
| 96 | Jamba 1.5 Large |
|
64.22 |
| 97 | Gemini 3 Pro |
|
64.21 |
| 98 | Qwen3 |
|
63.98 |
| 99 | Kimi K2.5 |
|
63.87 |
| 100 | Claude Sonnet 4 |
|
63.84 |