Quality
Highest overall scores
- 1
Qwen: Qwen3.8 Max58.1
- 2
Qwen: Qwen3.7 Plus39.4
- 3
Qwen: Qwen3.7 Flash0
Compare Alibaba models across answer quality, coding, complex tasks, speed, price, and context size.
Highest overall scores
Fastest output in this sample
Lowest price per 1M tokens
Overall capability · Higher is better
Output tokens per second · Higher is better
USD per 1M tokens · Lower is better
Quality benchmarks
Performance and price
3 models in the current catalogue.
| Model | Links | |||||
|---|---|---|---|---|---|---|
| Qwen: Qwen3.8 Max | 58.1 | 39 t/s | 1M | $2.00 | $6.00 | |
| Qwen: Qwen3.7 Plus | 39.4 | 15 t/s | 1M | $0.32 | $1.28 | |
| Qwen: Qwen3.7 Flash | 0 | 39 t/s | 1M | $0.03 | $0.13 |
Compare named evaluation records available for Alibaba models. Missing results remain blank rather than being estimated.
| Benchmark | Area | Qwen3.8 Max | Qwen3.7 Plus | Qwen3.7 Flash | Source |
|---|---|---|---|---|---|
| Intelligence Index | overallQuality | 58.1 | 39.4 | — | OpenRouter / artificial-analysis |
| Coding Index | coding | 71.8 | 55.9 | — | OpenRouter / artificial-analysis |
| Agentic Index | agentic | 58.4 | 20.7 | — | OpenRouter / artificial-analysis |
| GPQA | GPQA Diamond | 92.7 | 90 | — | OpenRouter model benchmarks |
| Humanity's Last Exam | HLE | 43 | 35.6 | — | OpenRouter model benchmarks |
| AA-LCR | AA-LCR | 74.3 | 69 | — | OpenRouter model benchmarks |
| GDPval-AA | GDPval-AA | 61.8 | 22.2 | — | OpenRouter model benchmarks |
| CritPt | CritPt | 20 | 9.1 | — | OpenRouter model benchmarks |
| SciCode | SciCode | 52.9 | 45.5 | — | OpenRouter model benchmarks |
| AA-Omniscience Accuracy | AA-Omniscience Accuracy | 31.9 | 22.5 | — | OpenRouter model benchmarks |
| AA-Omniscience Non-Hallucination Rate | AA-Omniscience Non-Hallucination Rate | 58.3 | 72.3 | — | OpenRouter model benchmarks |
| IFBench | IFBench | — | 78 | — | OpenRouter model benchmarks |
| τ²-Bench Telecom | τ²-Bench Telecom | — | 93 | — | OpenRouter model benchmarks |
| Terminal-Bench Hard | Terminal-Bench Hard | — | 47 | — | OpenRouter model benchmarks |
| Model | Quality | Evidence | Price /1k images | Links |
|---|---|---|---|---|
| Qwen Image 3 | — | — | — | View details |
| Qwen Image 3 Pro | — | — | — | View details |
| Model | Quality | Text to video | Image to video | Links |
|---|---|---|---|---|
| HappyHorse 1.1 | 88 | 1151 Elo | 1109 Elo | View details |
| HappyHorse 1.0 | 82 | 1132 Elo | 1091 Elo | View details |
| Wan 2.7 | 78 | 1109 Elo | 1092 Elo | View details |
| Wan 2.6 | 18 | 1029 Elo | 898 Elo | View details |