Quality
#14/7546.5
Overall quality score
See how Google: Gemini 3.1 Pro Preview compares with other broadly useful AI models across quality, speed, price, context, and practical capabilities.
46.5
Overall quality score
101 t/s
Output tokens per second
$2.00/1M
Per 1M tokens
$12.00/1M
Per 1M tokens
1M
Content read at once
Google: Gemini 3.1 Pro Preview ranks #14 of 75 for overall quality and #21 for response speed in the current catalogue.
Its listed input price is $2.00/1M, output price is $12.00/1M, and it can work with up to 1M of context at once.
Its strongest areas in the current data are following instructions, long documents, coding.
Named evaluation records for technical comparison. Results retain their original benchmark names and sources.
| Benchmark | Area | Gemini 3.1 Pro Preview | Source |
|---|---|---|---|
| Intelligence Index | overallQuality | 46.5 | OpenRouter model benchmarks |
| Coding Index | coding | 68.8 | OpenRouter model benchmarks |
| Agentic Index | agentic | 21.4 | OpenRouter model benchmarks |
| GPQA | GPQA Diamond | 94.1 | OpenRouter model benchmarks |
| Humanity's Last Exam | HLE | 44.7 | OpenRouter model benchmarks |
| IFBench | Instruction following | 77.1 | OpenRouter model benchmarks |
| τ²-Bench Telecom | Agentic tasks | 95.6 | OpenRouter model benchmarks |
| AA-LCR | Long-context reasoning | 72.7 | OpenRouter model benchmarks |
| GDPval-AA | Economically valuable tasks | 23.1 | OpenRouter model benchmarks |
| CritPt | Research-level physics | 17.7 | OpenRouter model benchmarks |
| SciCode | Scientific coding | 58.9 | OpenRouter model benchmarks |
| Terminal-Bench Hard | Agentic coding | 53.8 | OpenRouter model benchmarks |
| AA-Omniscience Accuracy | Factual knowledge | 55.2 | OpenRouter model benchmarks |
| AA-Omniscience Non-Hallucination Rate | Reliability | 50.1 | OpenRouter model benchmarks |
Quality benchmarks
The darker bar marks Google: Gemini 3.1 Pro Preview; the lighter bars provide context from other leading models.
Overall capability · Higher is better
Output tokens per second · Higher is better
USD per 1M tokens · Lower is better