Quality
Highest overall scores
- 1
Google: Gemini 3.5 Flash50.2
- 2
Google: Gemini 3.6 Flash50.1
- 3
Google: Gemini 3.1 Pro Preview46.5
Compare Google models across answer quality, coding, complex tasks, speed, price, and context size.
Highest overall scores
Fastest output in this sample
Lowest price per 1M tokens
Overall capability · Higher is better
Output tokens per second · Higher is better
USD per 1M tokens · Lower is better
Quality benchmarks
Performance and price
3 models in the current catalogue.
| Model | Links | |||||
|---|---|---|---|---|---|---|
| Google: Gemini 3.5 Flash | 50.2 | 119 t/s | 1M | $1.50 | $9.00 | |
| Google: Gemini 3.6 Flash | 50.1 | 148 t/s | 1M | $1.50 | $7.50 | |
| Google: Gemini 3.1 Pro Preview | 46.5 | 101 t/s | 1M | $2.00 | $12.00 |
Compare named evaluation records available for Google models. Missing results remain blank rather than being estimated.
| Benchmark | Area | Gemini 3.5 Flash | Gemini 3.6 Flash | Gemini 3.1 Pro Preview | Source |
|---|---|---|---|---|---|
| Intelligence Index | overallQuality | 50.2 | 50.1 | 46.5 | OpenRouter / artificial-analysis |
| Coding Index | coding | 70.1 | 69.2 | 68.8 | OpenRouter / artificial-analysis |
| Agentic Index | agentic | 37.4 | 38.7 | 21.4 | OpenRouter / artificial-analysis |
| GPQA | GPQA Diamond | — | 92.8 | 94.1 | OpenRouter model benchmarks |
| Humanity's Last Exam | HLE | — | 38.3 | 44.7 | OpenRouter model benchmarks |
| AA-LCR | AA-LCR | — | 69.7 | 72.7 | OpenRouter model benchmarks |
| GDPval-AA | GDPval-AA | — | 46.2 | 23.1 | OpenRouter model benchmarks |
| CritPt | CritPt | — | 10.6 | 17.7 | OpenRouter model benchmarks |
| SciCode | SciCode | — | 52.7 | 58.9 | OpenRouter model benchmarks |
| AA-Omniscience Accuracy | AA-Omniscience Accuracy | — | 50.2 | 55.2 | OpenRouter model benchmarks |
| AA-Omniscience Non-Hallucination Rate | AA-Omniscience Non-Hallucination Rate | — | 46.5 | 50.1 | OpenRouter model benchmarks |
| IFBench | Instruction following | — | — | 77.1 | OpenRouter model benchmarks |
| τ²-Bench Telecom | Agentic tasks | — | — | 95.6 | OpenRouter model benchmarks |
| Terminal-Bench Hard | Agentic coding | — | — | 53.8 | OpenRouter model benchmarks |
| Model | Quality | Evidence | Price /1k images | Links |
|---|---|---|---|---|
| Nano Banana 2 (Gemini 3.1 Flash Image Preview) | 1264 Elo | 19,895 samples | $67 | View details |
| Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) | 1263 Elo | 10,121 samples | $33.6 | View details |
| Nano Banana Pro (Gemini 3 Pro Image) | 1226 Elo | 12,280 samples | $134 | View details |
| Nano Banana (Gemini 2.5 Flash Image) | 1153 Elo | 11,198 samples | $39 | View details |
| Nano Banana 2 (Gemini 3.1 Flash Image) | — | — | — | View details |
| Nano Banana Pro (Gemini 3 Pro Image Preview) | — | — | — | View details |
| Model | Quality | Text to video | Image to video | Links |
|---|---|---|---|---|
| Veo 3.1 | 68 | 1099 Elo | 1086 Elo | View details |
| Veo 3.1 Fast | 58 | 1092 Elo | 1076 Elo | View details |
| Veo 3.1 Lite | 46 | 1090 Elo | 1064 Elo | View details |