Compare AI models.
Choose between two and five models to compare their quality, speed, context size, and pricing side by side.
Choose models.
Select up to five models.
2 of 5 selected
Side-by-side comparison.
Comparing 2 models.
| Model | Company | Links | ||||||
|---|---|---|---|---|---|---|---|---|
| Claude Opus 5 | 63.1 | 82 t/s | 1M | $0.018 | $5.00 | $25.00 | ||
| GPT-5.6 Sol | 60.9 | 44 t/s | 1.1M | $0.020 | $5.00 | $30.00 |
Cost per task estimates 1,000 input tokens and 500 output tokens using the listed API prices.
Named benchmark comparison.
Compare GPQA, MMLU-Pro, Humanity's Last Exam, SWE-bench, Terminal-Bench, AIME, and other available evaluation records directly.
| Benchmark | Area | Claude Opus 5 | GPT-5.6 Sol | Source |
|---|---|---|---|---|
| Intelligence Index | overallQuality | 63.1 | 60.9 | OpenRouter / artificial-analysis |
| Coding Index | coding | 78 | 77.4 | OpenRouter / artificial-analysis |
| Agentic Index | agentic | 59.2 | 57.8 | OpenRouter / artificial-analysis |
| GPQA | GPQA Diamond | 93.2 | 94.1 | OpenRouter model benchmarks |
| Humanity's Last Exam | HLE | 54.9 | 49.5 | OpenRouter model benchmarks |
| AA-LCR | AA-LCR | 75.7 | 77.7 | OpenRouter model benchmarks |
| GDPval-AA | GDPval-AA | 67.4 | 61.4 | OpenRouter model benchmarks |
| CritPt | CritPt | 29.1 | 32.3 | OpenRouter model benchmarks |
| SciCode | SciCode | 55.7 | 56.1 | OpenRouter model benchmarks |
| AA-Omniscience Accuracy | AA-Omniscience Accuracy | 60.9 | 59.4 | OpenRouter model benchmarks |
| AA-Omniscience Non-Hallucination Rate | AA-Omniscience Non-Hallucination Rate | 39.2 | 7.8 | OpenRouter model benchmarks |
| IFBench | IFBench | — | 72.7 | OpenRouter model benchmarks |
| τ²-Bench Telecom | τ²-Bench Telecom | — | 85.1 | OpenRouter model benchmarks |
| Terminal-Bench Hard | Terminal-Bench Hard | — | 65.9 | OpenRouter model benchmarks |