LLM Model Guide — 2026-08-25

Source: llm-stats.com leaderboard + OpenRouter pricing. Blend: 75% input / 25% output tokens. Speed ignored. Coverage: 330/358 priced, 26 unmatched.

Recommendations

The non-dominated frontier, ordered by weighted score (value × closeness to the frontier). Overpriced and far-off-frontier models drop out.

#ModelScore$/1MWeighted
1DeepSeek-V4-Pro-081354.5$0.54456.2
2DeepSeek-V4-Flash-073146.5$0.11346.9
3DeepSeek-V4-Flash-Vision-Exp48.7$0.33025.8
4GLM-5.354.7$2.15014.9
5Kimi K354.9$6.0005.6
6GPT-5.6 Sol57.4$11.255.1
7Claude Opus 556.2$10.004.5

Recommendation by task

TaskHighMediumLowFree
General / OverallGPT-5.6 Sol (57)DeepSeek-V4-Pro-0813 (54)DeepSeek-V4-Flash-0731 (46)GPT OSS 20B (18)
CodingGPT-5.6 Sol (51)DeepSeek-V4-Pro-0813 (44)DeepSeek-V4-Flash-0731 (36)IBM Granite 4.0 Tiny Preview (3)
MathQwen3.7 Max (44)DeepSeek-V4-Flash-Max (41)GPT OSS 20B High (32)GPT OSS 20B (18)
ReasoningGPT-5.6 Sol (57)DeepSeek-V4-Pro-0813 (52)DeepSeek-V4-Flash-0731 (44)GPT OSS 20B (16)
Tool Calling / AgentsKimi K3 (38)DeepSeek-V4-Flash-0731 (28)DeepSeek-V4-Flash-0731 (28)Llama 3.1 8B Instruct (17)
Search / RetrievalKimi K3 (35)Hy3 (25)Hy3 (25)Mistral NeMo Instruct (14)
Long ContextGemini 3.7 Flash (34)Gemini 3.7 Flash (34)Qwen3.5-9B (18)Nemotron 3.5 Lightning (30B A3B) (7)
Vision / MultimodalQwen3.8 Max (39)Qwen3.7-Plus (33)DeepSeek-V4-Flash-0423 (19)GPT OSS 120B (3)
Writing / CommunicationClaude Opus 4.6 (32)LongCat-Flash-Thinking-2601 (30)Qwen2.5 7B Instruct (14)Mistral Small 3 24B Instruct (8)
FinanceQwen3.7 Max (41)Qwen3.7-Plus (37)GPT OSS 120B (23)GPT OSS 20B (14)
LegalQwen3.7 Max (44)Qwen3.7-Plus (41)Qwen3.5-9B (24)GPT OSS 20B (14)
HealthcareQwen3.7 Max (44)Qwen3.7-Plus (41)GPT OSS 120B (32)GPT OSS 20B (16)

High = best score · Medium = best weighted (value × frontier-closeness) · Low = cheapest within 20 pts of the best · Free = lowest price.

Reasonable price curve (overall)

Fitted market price for a given overall score (log-linear fit across all priced models):

ScoreFair blended $/1M
0$0.327
10$0.483
20$0.713
30$1.052
40$1.553
50$2.292
60$3.382

(price doubles every ~17.8 points; n=284)

By category

For each category: the best models overall, the best within a ).

General / Overall

Recommended — DeepSeek-V4-Pro-0813 (score 54.5, $0.544/1M)

Best

#ModelScore$/1MPts/$
1GPT-5.6 Sol57.4$11.255
2Claude Opus 556.2$10.006
3Claude Fable 556.1$20.003
4Kimi K354.9$6.0009
5GLM-5.354.7$2.15025

Best under $2.000/1M

#ModelScore$/1MPts/$
1DeepSeek-V4-Pro-081354.5$0.544100
2Muse Spark 1.151.7$2.00026
3Gemini 3.7 Flash51.1$1.50034
4DeepSeek-V4-Flash-Vision-Exp48.7$0.330147
5Seed 2.1 Pro47.2$1.00047

Best value (score per $)

#ModelScore$/1MPts/$
1GPT OSS 20B High27.6$0.055502
2GPT OSS 120B29.6$0.070422
3DeepSeek-V4-Flash-073146.5$0.113413
4GPT OSS 120B High26.1$0.070371
5Laguna S 2.141.8$0.125334

Coding

Recommended — DeepSeek-V4-Pro-0813 (score 44.3, $0.544/1M)

Best

#ModelScore$/1MPts/$
1GPT-5.6 Sol50.6$11.255
2Claude Fable 548.8$20.002
3GPT-5.6 Terra46.4$4.50010
4Kimi K345.9$6.0008
5GLM-5.345.4$2.15021

Best under $2.000/1M

#ModelScore$/1MPts/$
1DeepSeek-V4-Pro-081344.3$0.54481
2GPT-5.6 Luna39.9$0.45089
3Gemini 3.7 Flash38.6$1.50026
4Qwen3.7 Max38.5$1.87521
5DeepSeek-V4-Flash-Vision-Exp38.5$0.330117

Best value (score per $)

#ModelScore$/1MPts/$
1DeepSeek-V4-Flash-073136.4$0.113324
2Laguna S 2.133.6$0.125269
3Muse Spark 1.233.6$0.125269
4DeepSeek-V4-Flash-042327.9$0.125223
5MiniCPM-SALA21.1$0.110192

Math

Recommended — DeepSeek-V4-Flash-Max (score 40.5, $0.175/1M)

Best

#ModelScore$/1MPts/$
1Qwen3.7 Max43.7$1.87523
2Claude Opus 543.3$10.004
3Gemini 3.1 Pro42.8$5.6258
4GLM-5.242.1$1.46229
5Claude Fable 541.9$20.002

Best under $2.000/1M

#ModelScore$/1MPts/$
1Qwen3.7 Max43.7$1.87523
2GLM-5.242.1$1.46229
3DeepSeek-V4-Pro-081341.8$0.54477
4DeepSeek-V4-Pro-Max41.4$2.00021
5Muse Spark 1.140.9$2.00020

Best value (score per $)

#ModelScore$/1MPts/$
1GPT OSS 20B High31.5$0.055573
2GPT OSS 120B24.6$0.070350
3GPT OSS 120B High24.2$0.070345
4GPT OSS 20B18.1$0.055329
5DeepSeek-V4-Flash-042338.0$0.125304

Reasoning

Recommended — DeepSeek-V4-Pro-0813 (score 52.0, $0.544/1M)

Best

#ModelScore$/1MPts/$
1GPT-5.6 Sol56.8$11.255
2Claude Opus 555.3$10.006
3GLM-5.354.9$2.15026
4Kimi K353.9$6.0009
5Claude Fable 553.5$20.003

Best under $2.000/1M

#ModelScore$/1MPts/$
1Muse Spark 1.152.3$2.00026
2DeepSeek-V4-Pro-081352.0$0.54496
3Gemini 3.7 Flash49.9$1.50033
4Seed 2.1 Pro48.1$1.00048
5Qwen3.7 Max47.1$1.87525

Best value (score per $)

#ModelScore$/1MPts/$
1GPT OSS 20B High27.7$0.055503
2DeepSeek-V4-Flash-073143.7$0.113388
3GPT OSS 120B High25.9$0.070369
4GPT OSS 120B23.8$0.070339
5Laguna S 2.141.9$0.125335

Tool Calling / Agents

Recommended — DeepSeek-V4-Flash-0731 (score 27.7, $0.113/1M)

Best

#ModelScore$/1MPts/$
1Kimi K338.3$6.0006
2GLM-5.336.5$2.15017
3GPT-5.6 Sol34.8$11.253
4Qwen3.8 Max34.6$3.00012
5DeepSeek-V4-Pro-081334.3$0.54463

Best under $2.000/1M

#ModelScore$/1MPts/$
1DeepSeek-V4-Pro-081334.3$0.54463
2Muse Spark 1.133.0$2.00016
3Llama 3.1 405B Instruct30.8$0.40077
4Gemini 3.7 Flash29.1$1.50019
5DeepSeek-V4-Flash-Vision-Exp28.3$0.33086

Best value (score per $)

#ModelScore$/1MPts/$
1Llama 3.1 8B Instruct17.0$0.057295
2DeepSeek-V4-Flash-073127.7$0.113246
3Muse Spark 1.221.2$0.125169
4Laguna S 2.119.2$0.125154
5DeepSeek-V4-Flash-042316.6$0.125133

Search / Retrieval

Recommended — Hy3 (score 25.2, $0.231/1M)

Best

#ModelScore$/1MPts/$
1Kimi K334.9$6.0006
2Claude Opus 530.7$10.003
3GPT-5.6 Sol29.5$11.253
4GPT-5.5 Pro28.5$67.500
5GPT-5.6 Terra27.6$4.5006

Best under $2.000/1M

#ModelScore$/1MPts/$
1Kimi K2.626.6$1.43819
2Seed 2.1 Pro25.3$1.00025
3Hy325.2$0.231109
4Seed 2.1 Turbo23.9$1.00024
5Kimi K2.522.5$1.20019

Best value (score per $)

#ModelScore$/1MPts/$
1Mistral NeMo Instruct13.8$0.022636
2Gemma 2 9B11.0$0.075147
3Hy325.2$0.231109
4Gemma 3n E4B8.1$0.075109
5DeepSeek-V4-Flash-Max14.1$0.17580

Long Context

Recommended — Gemini 3.7 Flash (score 33.6, $1.500/1M)

Best

#ModelScore$/1MPts/$
1Gemini 3.7 Flash33.6$1.50022
2GPT-5.6 Sol32.4$11.253
3Qwen3.8 Max32.0$3.00011
4Seed 2.1 Pro30.4$1.00030
5Seed 2.1 Turbo28.4$1.00028

Best under $2.000/1M

#ModelScore$/1MPts/$
1Gemini 3.7 Flash33.6$1.50022
2Seed 2.1 Pro30.4$1.00030
3Seed 2.1 Turbo28.4$1.00028
4Qwen3.6 Plus27.9$1.12525
5Qwen3.7-Plus27.6$0.56049

Best value (score per $)

#ModelScore$/1MPts/$
1Qwen3.5-9B18.1$0.112160
2Hy325.5$0.231111
3Gemma 4 12B16.2$0.160101
4Mistral Small 424.2$0.26292
5Nemotron 3 Super (120B A12B)14.9$0.16491

Vision / Multimodal

Recommended — Qwen3.7-Plus (score 33.3, $0.560/1M)

Best

#ModelScore$/1MPts/$
1Qwen3.8 Max38.9$3.00013
2Kimi K338.5$6.0006
3Claude Opus 538.5$10.004
4GPT-5.6 Sol37.8$11.253
5Claude Opus 4.837.7$10.004

Best under $2.000/1M

#ModelScore$/1MPts/$
1Muse Spark 1.136.1$2.00018
2Qwen3.8-27B35.3$1.05034
3Seed 2.1 Pro35.2$1.00035
4Gemini 3.7 Flash33.7$1.50022
5Qwen3.7-Plus33.3$0.56059

Best value (score per $)

#ModelScore$/1MPts/$
1DeepSeek-V4-Flash-042319.1$0.125153
2DeepSeek-V4-Flash-Max21.6$0.175124
3Step3-VL-10B17.5$0.150117
4MiMo-V2.523.9$0.210114
5Qwen3.5-9B12.1$0.112107

Writing / Communication

Recommended — LongCat-Flash-Thinking-2601 (score 30.0, $0.525/1M)

Best

#ModelScore$/1MPts/$
1Claude Opus 4.631.9$10.003
2LongCat-Flash-Thinking-260130.0$0.52557
3Claude Opus 4.527.2$10.003
4Claude Sonnet 4.627.1$6.0005
5Claude Sonnet 4.526.5$6.0004

Best under $2.000/1M

#ModelScore$/1MPts/$
1LongCat-Flash-Thinking-260130.0$0.52557
2Nova 2 Pro23.1$1.40017
3Nova 2 Omni23.1$0.85027
4Hermes 3 70B21.3$0.70030
5Nova 2 Lite20.4$0.85024

Best value (score per $)

#ModelScore$/1MPts/$
1Mistral Small 3 24B Instruct8.0$0.057139
2Qwen2.5 7B Instruct13.8$0.125110
3Llama-3.3 Nemotron Super 49B v117.3$0.164105
4Step3-VL-10B15.4$0.150103
5Qwen3.5-9B10.3$0.11291

Finance

Recommended — Qwen3.7-Plus (score 36.7, $0.560/1M)

Best

#ModelScore$/1MPts/$
1Qwen3.7 Max40.9$1.87522
2Qwen3.5-397B-A17B38.1$1.27530
3Qwen3.6 Plus36.8$1.12533
4Qwen3.7-Plus36.7$0.56066
5Sakana Namazu35.9$1.71221

Best under $2.000/1M

#ModelScore$/1MPts/$
1Qwen3.7 Max40.9$1.87522
2Qwen3.5-397B-A17B38.1$1.27530
3Qwen3.6 Plus36.8$1.12533
4Qwen3.7-Plus36.7$0.56066
5Sakana Namazu35.9$1.71221

Best value (score per $)

#ModelScore$/1MPts/$
1GPT OSS 120B23.3$0.070331
2GPT OSS 120B High18.4$0.070262
3GPT OSS 20B13.6$0.055246
4Nemotron 3.5 Lightning (30B A3B)21.1$0.088241
5DeepSeek-V4-Flash-042328.6$0.125229

Recommended — Qwen3.7-Plus (score 40.6, $0.560/1M)

Best

#ModelScore$/1MPts/$
1Qwen3.7 Max43.9$1.87523
2Qwen3.7-Plus40.6$0.56073
3Qwen3.6 Plus40.3$1.12536
4Qwen3.5-397B-A17B38.1$1.27530
5Sakana Namazu35.9$1.71221

Best under $2.000/1M

#ModelScore$/1MPts/$
1Qwen3.7 Max43.9$1.87523
2Qwen3.7-Plus40.6$0.56073
3Qwen3.6 Plus40.3$1.12536
4Qwen3.5-397B-A17B38.1$1.27530
5Sakana Namazu35.9$1.71221

Best value (score per $)

#ModelScore$/1MPts/$
1GPT OSS 120B23.3$0.070331
2GPT OSS 120B High18.4$0.070262
3GPT OSS 20B13.6$0.055246
4Nemotron 3.5 Lightning (30B A3B)21.1$0.088241
5DeepSeek-V4-Flash-042328.6$0.125229

Healthcare

Recommended — Qwen3.7-Plus (score 41.4, $0.560/1M)

Best

#ModelScore$/1MPts/$
1Qwen3.7 Max44.2$1.87524
2Qwen3.7-Plus41.4$0.56074
3Qwen3.6 Plus41.3$1.12537
4Qwen3.5-397B-A17B38.5$1.27530
5Qwen3.5-122B-A10B37.8$0.71553

Best under $2.000/1M

#ModelScore$/1MPts/$
1Qwen3.7 Max44.2$1.87524
2Qwen3.7-Plus41.4$0.56074
3Qwen3.6 Plus41.3$1.12537
4Qwen3.5-397B-A17B38.5$1.27530
5Qwen3.5-122B-A10B37.8$0.71553

Best value (score per $)

#ModelScore$/1MPts/$
1GPT OSS 120B31.6$0.070450
2GPT OSS 20B16.2$0.055294
3GPT OSS 120B High18.4$0.070262
4Nemotron 3.5 Lightning (30B A3B)21.1$0.088241
5DeepSeek-V4-Flash-042328.6$0.125229

Data notes

  • Scores are llm-stats index_* composites (0–~60 scale, relative to the 2026 field).
  • Prices are USD per 1M tokens; $/1M is a blended input/output cost. Open-weight models priced via OpenRouter top provider; frontier models use official API price.
  • Value = score ÷ blended price; Weighted = value × e^-(gap to the frontier) — rewards cheap-but-near-SOTA, penalises far-from-SOTA and expensive-SOTA.
  • The ‘reasonable price’ curve is a market fit, not a recommendation to pay that much.