LLM Model Guide — 2026-08-25
Source: llm-stats.com leaderboard + OpenRouter pricing. Blend: 75% input / 25% output tokens. Speed ignored. Coverage: 330/358 priced, 26 unmatched.
Recommendations
The non-dominated frontier, ordered by weighted score (value × closeness to the frontier). Overpriced and far-off-frontier models drop out.
| # | Model | Score | $/1M | Weighted |
|---|---|---|---|---|
| 1 | DeepSeek-V4-Pro-0813 | 54.5 | $0.544 | 56.2 |
| 2 | DeepSeek-V4-Flash-0731 | 46.5 | $0.113 | 46.9 |
| 3 | DeepSeek-V4-Flash-Vision-Exp | 48.7 | $0.330 | 25.8 |
| 4 | GLM-5.3 | 54.7 | $2.150 | 14.9 |
| 5 | Kimi K3 | 54.9 | $6.000 | 5.6 |
| 6 | GPT-5.6 Sol | 57.4 | $11.25 | 5.1 |
| 7 | Claude Opus 5 | 56.2 | $10.00 | 4.5 |
Recommendation by task
| Task | High | Medium | Low | Free |
|---|---|---|---|---|
| General / Overall | GPT-5.6 Sol (57) | DeepSeek-V4-Pro-0813 (54) | DeepSeek-V4-Flash-0731 (46) | GPT OSS 20B (18) |
| Coding | GPT-5.6 Sol (51) | DeepSeek-V4-Pro-0813 (44) | DeepSeek-V4-Flash-0731 (36) | IBM Granite 4.0 Tiny Preview (3) |
| Math | Qwen3.7 Max (44) | DeepSeek-V4-Flash-Max (41) | GPT OSS 20B High (32) | GPT OSS 20B (18) |
| Reasoning | GPT-5.6 Sol (57) | DeepSeek-V4-Pro-0813 (52) | DeepSeek-V4-Flash-0731 (44) | GPT OSS 20B (16) |
| Tool Calling / Agents | Kimi K3 (38) | DeepSeek-V4-Flash-0731 (28) | DeepSeek-V4-Flash-0731 (28) | Llama 3.1 8B Instruct (17) |
| Search / Retrieval | Kimi K3 (35) | Hy3 (25) | Hy3 (25) | Mistral NeMo Instruct (14) |
| Long Context | Gemini 3.7 Flash (34) | Gemini 3.7 Flash (34) | Qwen3.5-9B (18) | Nemotron 3.5 Lightning (30B A3B) (7) |
| Vision / Multimodal | Qwen3.8 Max (39) | Qwen3.7-Plus (33) | DeepSeek-V4-Flash-0423 (19) | GPT OSS 120B (3) |
| Writing / Communication | Claude Opus 4.6 (32) | LongCat-Flash-Thinking-2601 (30) | Qwen2.5 7B Instruct (14) | Mistral Small 3 24B Instruct (8) |
| Finance | Qwen3.7 Max (41) | Qwen3.7-Plus (37) | GPT OSS 120B (23) | GPT OSS 20B (14) |
| Legal | Qwen3.7 Max (44) | Qwen3.7-Plus (41) | Qwen3.5-9B (24) | GPT OSS 20B (14) |
| Healthcare | Qwen3.7 Max (44) | Qwen3.7-Plus (41) | GPT OSS 120B (32) | GPT OSS 20B (16) |
High = best score · Medium = best weighted (value × frontier-closeness) · Low = cheapest within 20 pts of the best · Free = lowest price.
Reasonable price curve (overall)
Fitted market price for a given overall score (log-linear fit across all priced models):
| Score | Fair blended $/1M |
|---|---|
| 0 | $0.327 |
| 10 | $0.483 |
| 20 | $0.713 |
| 30 | $1.052 |
| 40 | $1.553 |
| 50 | $2.292 |
| 60 | $3.382 |
(price doubles every ~17.8 points; n=284)
By category
For each category: the best models overall, the best within a ).
General / Overall
Recommended — DeepSeek-V4-Pro-0813 (score 54.5, $0.544/1M)
Best
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | GPT-5.6 Sol | 57.4 | $11.25 | 5 |
| 2 | Claude Opus 5 | 56.2 | $10.00 | 6 |
| 3 | Claude Fable 5 | 56.1 | $20.00 | 3 |
| 4 | Kimi K3 | 54.9 | $6.000 | 9 |
| 5 | GLM-5.3 | 54.7 | $2.150 | 25 |
Best under $2.000/1M
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | DeepSeek-V4-Pro-0813 | 54.5 | $0.544 | 100 |
| 2 | Muse Spark 1.1 | 51.7 | $2.000 | 26 |
| 3 | Gemini 3.7 Flash | 51.1 | $1.500 | 34 |
| 4 | DeepSeek-V4-Flash-Vision-Exp | 48.7 | $0.330 | 147 |
| 5 | Seed 2.1 Pro | 47.2 | $1.000 | 47 |
Best value (score per $)
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | GPT OSS 20B High | 27.6 | $0.055 | 502 |
| 2 | GPT OSS 120B | 29.6 | $0.070 | 422 |
| 3 | DeepSeek-V4-Flash-0731 | 46.5 | $0.113 | 413 |
| 4 | GPT OSS 120B High | 26.1 | $0.070 | 371 |
| 5 | Laguna S 2.1 | 41.8 | $0.125 | 334 |
Coding
Recommended — DeepSeek-V4-Pro-0813 (score 44.3, $0.544/1M)
Best
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | GPT-5.6 Sol | 50.6 | $11.25 | 5 |
| 2 | Claude Fable 5 | 48.8 | $20.00 | 2 |
| 3 | GPT-5.6 Terra | 46.4 | $4.500 | 10 |
| 4 | Kimi K3 | 45.9 | $6.000 | 8 |
| 5 | GLM-5.3 | 45.4 | $2.150 | 21 |
Best under $2.000/1M
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | DeepSeek-V4-Pro-0813 | 44.3 | $0.544 | 81 |
| 2 | GPT-5.6 Luna | 39.9 | $0.450 | 89 |
| 3 | Gemini 3.7 Flash | 38.6 | $1.500 | 26 |
| 4 | Qwen3.7 Max | 38.5 | $1.875 | 21 |
| 5 | DeepSeek-V4-Flash-Vision-Exp | 38.5 | $0.330 | 117 |
Best value (score per $)
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | DeepSeek-V4-Flash-0731 | 36.4 | $0.113 | 324 |
| 2 | Laguna S 2.1 | 33.6 | $0.125 | 269 |
| 3 | Muse Spark 1.2 | 33.6 | $0.125 | 269 |
| 4 | DeepSeek-V4-Flash-0423 | 27.9 | $0.125 | 223 |
| 5 | MiniCPM-SALA | 21.1 | $0.110 | 192 |
Math
Recommended — DeepSeek-V4-Flash-Max (score 40.5, $0.175/1M)
Best
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | Qwen3.7 Max | 43.7 | $1.875 | 23 |
| 2 | Claude Opus 5 | 43.3 | $10.00 | 4 |
| 3 | Gemini 3.1 Pro | 42.8 | $5.625 | 8 |
| 4 | GLM-5.2 | 42.1 | $1.462 | 29 |
| 5 | Claude Fable 5 | 41.9 | $20.00 | 2 |
Best under $2.000/1M
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | Qwen3.7 Max | 43.7 | $1.875 | 23 |
| 2 | GLM-5.2 | 42.1 | $1.462 | 29 |
| 3 | DeepSeek-V4-Pro-0813 | 41.8 | $0.544 | 77 |
| 4 | DeepSeek-V4-Pro-Max | 41.4 | $2.000 | 21 |
| 5 | Muse Spark 1.1 | 40.9 | $2.000 | 20 |
Best value (score per $)
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | GPT OSS 20B High | 31.5 | $0.055 | 573 |
| 2 | GPT OSS 120B | 24.6 | $0.070 | 350 |
| 3 | GPT OSS 120B High | 24.2 | $0.070 | 345 |
| 4 | GPT OSS 20B | 18.1 | $0.055 | 329 |
| 5 | DeepSeek-V4-Flash-0423 | 38.0 | $0.125 | 304 |
Reasoning
Recommended — DeepSeek-V4-Pro-0813 (score 52.0, $0.544/1M)
Best
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | GPT-5.6 Sol | 56.8 | $11.25 | 5 |
| 2 | Claude Opus 5 | 55.3 | $10.00 | 6 |
| 3 | GLM-5.3 | 54.9 | $2.150 | 26 |
| 4 | Kimi K3 | 53.9 | $6.000 | 9 |
| 5 | Claude Fable 5 | 53.5 | $20.00 | 3 |
Best under $2.000/1M
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | Muse Spark 1.1 | 52.3 | $2.000 | 26 |
| 2 | DeepSeek-V4-Pro-0813 | 52.0 | $0.544 | 96 |
| 3 | Gemini 3.7 Flash | 49.9 | $1.500 | 33 |
| 4 | Seed 2.1 Pro | 48.1 | $1.000 | 48 |
| 5 | Qwen3.7 Max | 47.1 | $1.875 | 25 |
Best value (score per $)
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | GPT OSS 20B High | 27.7 | $0.055 | 503 |
| 2 | DeepSeek-V4-Flash-0731 | 43.7 | $0.113 | 388 |
| 3 | GPT OSS 120B High | 25.9 | $0.070 | 369 |
| 4 | GPT OSS 120B | 23.8 | $0.070 | 339 |
| 5 | Laguna S 2.1 | 41.9 | $0.125 | 335 |
Tool Calling / Agents
Recommended — DeepSeek-V4-Flash-0731 (score 27.7, $0.113/1M)
Best
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | Kimi K3 | 38.3 | $6.000 | 6 |
| 2 | GLM-5.3 | 36.5 | $2.150 | 17 |
| 3 | GPT-5.6 Sol | 34.8 | $11.25 | 3 |
| 4 | Qwen3.8 Max | 34.6 | $3.000 | 12 |
| 5 | DeepSeek-V4-Pro-0813 | 34.3 | $0.544 | 63 |
Best under $2.000/1M
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | DeepSeek-V4-Pro-0813 | 34.3 | $0.544 | 63 |
| 2 | Muse Spark 1.1 | 33.0 | $2.000 | 16 |
| 3 | Llama 3.1 405B Instruct | 30.8 | $0.400 | 77 |
| 4 | Gemini 3.7 Flash | 29.1 | $1.500 | 19 |
| 5 | DeepSeek-V4-Flash-Vision-Exp | 28.3 | $0.330 | 86 |
Best value (score per $)
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | Llama 3.1 8B Instruct | 17.0 | $0.057 | 295 |
| 2 | DeepSeek-V4-Flash-0731 | 27.7 | $0.113 | 246 |
| 3 | Muse Spark 1.2 | 21.2 | $0.125 | 169 |
| 4 | Laguna S 2.1 | 19.2 | $0.125 | 154 |
| 5 | DeepSeek-V4-Flash-0423 | 16.6 | $0.125 | 133 |
Search / Retrieval
Recommended — Hy3 (score 25.2, $0.231/1M)
Best
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | Kimi K3 | 34.9 | $6.000 | 6 |
| 2 | Claude Opus 5 | 30.7 | $10.00 | 3 |
| 3 | GPT-5.6 Sol | 29.5 | $11.25 | 3 |
| 4 | GPT-5.5 Pro | 28.5 | $67.50 | 0 |
| 5 | GPT-5.6 Terra | 27.6 | $4.500 | 6 |
Best under $2.000/1M
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | Kimi K2.6 | 26.6 | $1.438 | 19 |
| 2 | Seed 2.1 Pro | 25.3 | $1.000 | 25 |
| 3 | Hy3 | 25.2 | $0.231 | 109 |
| 4 | Seed 2.1 Turbo | 23.9 | $1.000 | 24 |
| 5 | Kimi K2.5 | 22.5 | $1.200 | 19 |
Best value (score per $)
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | Mistral NeMo Instruct | 13.8 | $0.022 | 636 |
| 2 | Gemma 2 9B | 11.0 | $0.075 | 147 |
| 3 | Hy3 | 25.2 | $0.231 | 109 |
| 4 | Gemma 3n E4B | 8.1 | $0.075 | 109 |
| 5 | DeepSeek-V4-Flash-Max | 14.1 | $0.175 | 80 |
Long Context
Recommended — Gemini 3.7 Flash (score 33.6, $1.500/1M)
Best
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | Gemini 3.7 Flash | 33.6 | $1.500 | 22 |
| 2 | GPT-5.6 Sol | 32.4 | $11.25 | 3 |
| 3 | Qwen3.8 Max | 32.0 | $3.000 | 11 |
| 4 | Seed 2.1 Pro | 30.4 | $1.000 | 30 |
| 5 | Seed 2.1 Turbo | 28.4 | $1.000 | 28 |
Best under $2.000/1M
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | Gemini 3.7 Flash | 33.6 | $1.500 | 22 |
| 2 | Seed 2.1 Pro | 30.4 | $1.000 | 30 |
| 3 | Seed 2.1 Turbo | 28.4 | $1.000 | 28 |
| 4 | Qwen3.6 Plus | 27.9 | $1.125 | 25 |
| 5 | Qwen3.7-Plus | 27.6 | $0.560 | 49 |
Best value (score per $)
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | Qwen3.5-9B | 18.1 | $0.112 | 160 |
| 2 | Hy3 | 25.5 | $0.231 | 111 |
| 3 | Gemma 4 12B | 16.2 | $0.160 | 101 |
| 4 | Mistral Small 4 | 24.2 | $0.262 | 92 |
| 5 | Nemotron 3 Super (120B A12B) | 14.9 | $0.164 | 91 |
Vision / Multimodal
Recommended — Qwen3.7-Plus (score 33.3, $0.560/1M)
Best
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | Qwen3.8 Max | 38.9 | $3.000 | 13 |
| 2 | Kimi K3 | 38.5 | $6.000 | 6 |
| 3 | Claude Opus 5 | 38.5 | $10.00 | 4 |
| 4 | GPT-5.6 Sol | 37.8 | $11.25 | 3 |
| 5 | Claude Opus 4.8 | 37.7 | $10.00 | 4 |
Best under $2.000/1M
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | Muse Spark 1.1 | 36.1 | $2.000 | 18 |
| 2 | Qwen3.8-27B | 35.3 | $1.050 | 34 |
| 3 | Seed 2.1 Pro | 35.2 | $1.000 | 35 |
| 4 | Gemini 3.7 Flash | 33.7 | $1.500 | 22 |
| 5 | Qwen3.7-Plus | 33.3 | $0.560 | 59 |
Best value (score per $)
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | DeepSeek-V4-Flash-0423 | 19.1 | $0.125 | 153 |
| 2 | DeepSeek-V4-Flash-Max | 21.6 | $0.175 | 124 |
| 3 | Step3-VL-10B | 17.5 | $0.150 | 117 |
| 4 | MiMo-V2.5 | 23.9 | $0.210 | 114 |
| 5 | Qwen3.5-9B | 12.1 | $0.112 | 107 |
Writing / Communication
Recommended — LongCat-Flash-Thinking-2601 (score 30.0, $0.525/1M)
Best
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | Claude Opus 4.6 | 31.9 | $10.00 | 3 |
| 2 | LongCat-Flash-Thinking-2601 | 30.0 | $0.525 | 57 |
| 3 | Claude Opus 4.5 | 27.2 | $10.00 | 3 |
| 4 | Claude Sonnet 4.6 | 27.1 | $6.000 | 5 |
| 5 | Claude Sonnet 4.5 | 26.5 | $6.000 | 4 |
Best under $2.000/1M
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | LongCat-Flash-Thinking-2601 | 30.0 | $0.525 | 57 |
| 2 | Nova 2 Pro | 23.1 | $1.400 | 17 |
| 3 | Nova 2 Omni | 23.1 | $0.850 | 27 |
| 4 | Hermes 3 70B | 21.3 | $0.700 | 30 |
| 5 | Nova 2 Lite | 20.4 | $0.850 | 24 |
Best value (score per $)
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | Mistral Small 3 24B Instruct | 8.0 | $0.057 | 139 |
| 2 | Qwen2.5 7B Instruct | 13.8 | $0.125 | 110 |
| 3 | Llama-3.3 Nemotron Super 49B v1 | 17.3 | $0.164 | 105 |
| 4 | Step3-VL-10B | 15.4 | $0.150 | 103 |
| 5 | Qwen3.5-9B | 10.3 | $0.112 | 91 |
Finance
Recommended — Qwen3.7-Plus (score 36.7, $0.560/1M)
Best
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | Qwen3.7 Max | 40.9 | $1.875 | 22 |
| 2 | Qwen3.5-397B-A17B | 38.1 | $1.275 | 30 |
| 3 | Qwen3.6 Plus | 36.8 | $1.125 | 33 |
| 4 | Qwen3.7-Plus | 36.7 | $0.560 | 66 |
| 5 | Sakana Namazu | 35.9 | $1.712 | 21 |
Best under $2.000/1M
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | Qwen3.7 Max | 40.9 | $1.875 | 22 |
| 2 | Qwen3.5-397B-A17B | 38.1 | $1.275 | 30 |
| 3 | Qwen3.6 Plus | 36.8 | $1.125 | 33 |
| 4 | Qwen3.7-Plus | 36.7 | $0.560 | 66 |
| 5 | Sakana Namazu | 35.9 | $1.712 | 21 |
Best value (score per $)
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | GPT OSS 120B | 23.3 | $0.070 | 331 |
| 2 | GPT OSS 120B High | 18.4 | $0.070 | 262 |
| 3 | GPT OSS 20B | 13.6 | $0.055 | 246 |
| 4 | Nemotron 3.5 Lightning (30B A3B) | 21.1 | $0.088 | 241 |
| 5 | DeepSeek-V4-Flash-0423 | 28.6 | $0.125 | 229 |
Legal
Recommended — Qwen3.7-Plus (score 40.6, $0.560/1M)
Best
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | Qwen3.7 Max | 43.9 | $1.875 | 23 |
| 2 | Qwen3.7-Plus | 40.6 | $0.560 | 73 |
| 3 | Qwen3.6 Plus | 40.3 | $1.125 | 36 |
| 4 | Qwen3.5-397B-A17B | 38.1 | $1.275 | 30 |
| 5 | Sakana Namazu | 35.9 | $1.712 | 21 |
Best under $2.000/1M
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | Qwen3.7 Max | 43.9 | $1.875 | 23 |
| 2 | Qwen3.7-Plus | 40.6 | $0.560 | 73 |
| 3 | Qwen3.6 Plus | 40.3 | $1.125 | 36 |
| 4 | Qwen3.5-397B-A17B | 38.1 | $1.275 | 30 |
| 5 | Sakana Namazu | 35.9 | $1.712 | 21 |
Best value (score per $)
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | GPT OSS 120B | 23.3 | $0.070 | 331 |
| 2 | GPT OSS 120B High | 18.4 | $0.070 | 262 |
| 3 | GPT OSS 20B | 13.6 | $0.055 | 246 |
| 4 | Nemotron 3.5 Lightning (30B A3B) | 21.1 | $0.088 | 241 |
| 5 | DeepSeek-V4-Flash-0423 | 28.6 | $0.125 | 229 |
Healthcare
Recommended — Qwen3.7-Plus (score 41.4, $0.560/1M)
Best
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | Qwen3.7 Max | 44.2 | $1.875 | 24 |
| 2 | Qwen3.7-Plus | 41.4 | $0.560 | 74 |
| 3 | Qwen3.6 Plus | 41.3 | $1.125 | 37 |
| 4 | Qwen3.5-397B-A17B | 38.5 | $1.275 | 30 |
| 5 | Qwen3.5-122B-A10B | 37.8 | $0.715 | 53 |
Best under $2.000/1M
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | Qwen3.7 Max | 44.2 | $1.875 | 24 |
| 2 | Qwen3.7-Plus | 41.4 | $0.560 | 74 |
| 3 | Qwen3.6 Plus | 41.3 | $1.125 | 37 |
| 4 | Qwen3.5-397B-A17B | 38.5 | $1.275 | 30 |
| 5 | Qwen3.5-122B-A10B | 37.8 | $0.715 | 53 |
Best value (score per $)
| # | Model | Score | $/1M | Pts/$ |
|---|---|---|---|---|
| 1 | GPT OSS 120B | 31.6 | $0.070 | 450 |
| 2 | GPT OSS 20B | 16.2 | $0.055 | 294 |
| 3 | GPT OSS 120B High | 18.4 | $0.070 | 262 |
| 4 | Nemotron 3.5 Lightning (30B A3B) | 21.1 | $0.088 | 241 |
| 5 | DeepSeek-V4-Flash-0423 | 28.6 | $0.125 | 229 |
Data notes
- Scores are llm-stats
index_*composites (0–~60 scale, relative to the 2026 field). - Prices are USD per 1M tokens;
$/1Mis a blended input/output cost. Open-weight models priced via OpenRouter top provider; frontier models use official API price. - Value = score ÷ blended price; Weighted = value × e^-(gap to the frontier) — rewards cheap-but-near-SOTA, penalises far-from-SOTA and expensive-SOTA.
- The ‘reasonable price’ curve is a market fit, not a recommendation to pay that much.