Quality
Highest overall scores
- 1
Mistral: Mistral Medium 3.530.4
- 2
Mistral: Mistral Small 419.7
- 3
Mistral: Mistral Large 3 251215.9
Compare Mistral AI models across answer quality, coding, complex tasks, speed, price, and context size.
Highest overall scores
Fastest output in this sample
Lowest price per 1M tokens
Overall capability · Higher is better
Output tokens per second · Higher is better
USD per 1M tokens · Lower is better
Quality benchmarks
Performance and price
3 models in the current catalogue.
| Model | Links | |||||
|---|---|---|---|---|---|---|
| Mistral: Mistral Medium 3.5 | 30.4 | 33 t/s | 262.1K | $1.50 | $7.50 | |
| Mistral: Mistral Small 4 | 19.7 | 92 t/s | 262.1K | $0.15 | $0.60 | |
| Mistral: Mistral Large 3 2512 | 15.9 | 40 t/s | 262.1K | $0.50 | $1.50 |
Compare named evaluation records available for Mistral AI models. Missing results remain blank rather than being estimated.
| Benchmark | Area | Mistral Medium 3.5 | Mistral Small 4 | Mistral Large 3 2512 | Source |
|---|---|---|---|---|---|
| Intelligence Index | overallQuality | 30.4 | 19.7 | 15.9 | OpenRouter / artificial-analysis |
| Coding Index | coding | 46.9 | 26.6 | 20.1 | OpenRouter / artificial-analysis |
| Agentic Index | agentic | 19.2 | 4.6 | 5.5 | OpenRouter / artificial-analysis |
| GPQA | GPQA Diamond | 74.8 | 76.9 | 68 | OpenRouter model benchmarks |
| Humanity's Last Exam | HLE | 13.8 | 9.9 | 4.2 | OpenRouter model benchmarks |
| IFBench | IFBench | 68.8 | 48.2 | 36.2 | OpenRouter model benchmarks |
| τ²-Bench Telecom | τ²-Bench Telecom | 94.2 | 41.2 | 24.6 | OpenRouter model benchmarks |
| AA-LCR | AA-LCR | 65.3 | 47.3 | 34.7 | OpenRouter model benchmarks |
| GDPval-AA | GDPval-AA | 21.6 | 4.5 | 7 | OpenRouter model benchmarks |
| CritPt | CritPt | 0 | 0.3 | 0 | OpenRouter model benchmarks |
| SciCode | SciCode | 39.6 | 38 | 36.2 | OpenRouter model benchmarks |
| Terminal-Bench Hard | Terminal-Bench Hard | 33.3 | 17.4 | 15.9 | OpenRouter model benchmarks |
| AA-Omniscience Accuracy | AA-Omniscience Accuracy | 24.7 | 21.7 | 24.9 | OpenRouter model benchmarks |
| AA-Omniscience Non-Hallucination Rate | AA-Omniscience Non-Hallucination Rate | 18.4 | 33.5 | 14 | OpenRouter model benchmarks |