Find the right AI model
What will you use it for?
Choose the main type of task you plan to do
Choose a use case to get evidence-based recommendations. How recommendations work
| Model | Company | Links | |||||||
|---|---|---|---|---|---|---|---|---|---|
| 1 | 1.Claude Opus 5 | 60.7 | 63 | 1M | $0.018 | $5.00 | $25.00 | ||
| 2 | 2.GPT-5.6 Sol | 58.9 | 85 | 1.1M | $0.020 | $5.00 | $30.00 | ||
| 3 | 3.Kimi K3 | 57.1 | 73 | 1M | $0.011 | $3.00 | $15.00 | ||
| 4 | 4.Claude Opus 4.8 | 55.7 | Not measured | 1M | $0.018 | $5.00 | $25.00 | ||
| 5 | 5.GPT-5.6 Terra | 55 | Not measured | 1.1M | $0.0040 | $1.00 | $6.00 | ||
| 6 | 6.Grok 4.5 | 53.8 | 54 | 500K | $0.0050 | $2.00 | $6.00 | ||
| 7 | 7.Claude Sonnet 5 | 53.4 | 96.5 | 1M | $0.0070 | $2.00 | $10.00 | ||
| 8 | 8.GPT-5.6 Luna | 51.2 | Not measured | 1.1M | $0.0004 | $0.10 | $0.60 | ||
| 9 | 9.GLM 5.2 | 51.1 | 159 | 1M | $0.0029 | $1.12 | $3.52 | ||
| 10 | 10.Muse Spark 1.1 | 50.6 | 143 | 1M | $0.0034 | $1.25 | $4.25 | ||
| 11 | 11.Gemini 3.5 Flash | 50.2 | 119 | 1M | $0.0060 | $1.50 | $9.00 | ||
| 12 | 12.Gemini 3.6 Flash | 50.1 | 148 | 1M | $0.0053 | $1.50 | $7.50 | ||
| 13 | 13.DeepSeek V4 Flash 0731 | 49.9 | Not measured | 1M | $0.0003 | $0.14 | $0.28 | ||
| 14 | 14.Gemini 3.1 Pro Preview | 46.5 | 101 | 1M | $0.0080 | $2.00 | $12.00 | ||
| 15 | 15.Qwen3.7 Max | 46 | 35 | 1M | $0.0037 | $1.48 | $4.42 | ||
| 16 | 16.MiniMax M3 | 44.4 | 83 | 1M | $0.0009 | $0.30 | $1.20 | ||
| 17 | 17.DeepSeek V4 Pro | 44.3 | 87 | 1M | $0.0009 | $0.43 | $0.87 | ||
| 18 | 18.MiMo-V2.5-Pro | 42.2 | 68 | 1.1M | $0.0009 | $0.43 | $0.87 | ||
| 19 | 19.Kimi K2.7 Code | 41.9 | 168 | 262.1K | $0.0025 | $0.73 | $3.50 | ||
| 20 | 20.Hy3 preview | 41.2 | 27 | 262.1K | $0.0002 | $0.06 | $0.21 |
Showing 1β20 of 75
1 / 4
Cost per task estimates 1,000 input tokens and 500 output tokens using the listed API prices. Recommendation scores include an uncertainty penalty when relevant benchmark evidence is incomplete.