Anthropic model benchmarks.

Compare Anthropic models across answer quality, coding, complex tasks, speed, price, and context size.

Quality

Highest overall scores

  1. 1Claude Opus 563.1
  2. 2Anthropic: Claude Opus 4.857.3
  3. 3Anthropic: Claude Sonnet 555.3

Response speed

Fastest output in this sample

  1. 1Anthropic: Claude Opus 4.8 (Fast)148.5 t/s
  2. 2Claude Opus 5 (Fast)95 t/s
  3. 3Claude Opus 582 t/s

Input price

Lowest price per 1M tokens

  1. 1Anthropic: Claude Sonnet 5$2.00/1M
  2. 2Claude Opus 5$5.00/1M
  3. 3Anthropic: Claude Opus 4.8$5.00/1M

Highlights

Quality

Overall capability · Higher is better

Speed

Output tokens per second · Higher is better

Input price

USD per 1M tokens · Lower is better

Quality benchmarks

Compare Anthropic model capability.

Performance and price

Compare practical tradeoffs.

All Anthropic models.

5 models in the current catalogue.

Anthropic technical benchmarks.

Compare named evaluation records available for Anthropic models. Missing results remain blank rather than being estimated.

BenchmarkAreaClaude Opus 5Claude Opus 4.8Claude Sonnet 5Claude Opus 4.8 (Fast)Claude Opus 5 (Fast)Source
Intelligence IndexoverallQuality63.157.355.3——OpenRouter / artificial-analysis
Coding Indexcoding7874.371.5——OpenRouter / artificial-analysis
Agentic Indexagentic59.249.449.7——OpenRouter / artificial-analysis
GPQAGPQA Diamond93.29291.1——OpenRouter model benchmarks
Humanity's Last ExamHLE54.948.741.3——OpenRouter model benchmarks
AA-LCRAA-LCR75.77377——OpenRouter model benchmarks
GDPval-AAGDPval-AA67.454.354.9——OpenRouter model benchmarks
CritPtCritPt29.120.916.9——OpenRouter model benchmarks
SciCodeSciCode55.753.553.6——OpenRouter model benchmarks
AA-Omniscience AccuracyAA-Omniscience Accuracy60.948.840.1——OpenRouter model benchmarks
AA-Omniscience Non-Hallucination RateAA-Omniscience Non-Hallucination Rate39.260.760.6——OpenRouter model benchmarks
IFBenchIFBench—62.2———OpenRouter model benchmarks
τ²-Bench Telecomτ²-Bench Telecom—94.4———OpenRouter model benchmarks
Terminal-Bench HardTerminal-Bench Hard—58.3———OpenRouter model benchmarks