Anthropic model benchmarks.

Compare Anthropic models across answer quality, coding, complex tasks, speed, price, and context size.

Quality

Highest overall scores

  1. 1Claude Opus 560.7
  2. 2Anthropic: Claude Opus 4.855.7
  3. 3Anthropic: Claude Sonnet 553.4

Response speed

Fastest output in this sample

  1. 1Claude Opus 5 (Fast)112 t/s
  2. 2Anthropic: Claude Sonnet 596.5 t/s
  3. 3Claude Opus 563 t/s

Input price

Lowest price per 1M tokens

  1. 1Anthropic: Claude Sonnet 5$2.00/1M
  2. 2Claude Opus 5$5.00/1M
  3. 3Anthropic: Claude Opus 4.8$5.00/1M

Highlights

Quality

Overall capability · Higher is better

Speed

Output tokens per second · Higher is better

Input price

USD per 1M tokens · Lower is better

Quality benchmarks

Compare Anthropic model capability.

Performance and price

Compare practical tradeoffs.

All Anthropic models.

5 models in the current catalogue.

Anthropic technical benchmarks.

Compare named evaluation records available for Anthropic models. Missing results remain blank rather than being estimated.

BenchmarkAreaClaude Opus 5Claude Opus 4.8Claude Sonnet 5Claude Opus 4.8 (Fast)Claude Opus 5 (Fast)Source
Intelligence IndexoverallQuality60.755.753.4OpenRouter / artificial-analysis
Coding Indexcoding7874.371.5OpenRouter / artificial-analysis
Agentic Indexagentic55.347.246.7OpenRouter / artificial-analysis
GPQAGPQA Diamond93.29291.1OpenRouter model benchmarks
Humanity's Last ExamHLE52.645.739.6OpenRouter model benchmarks
AA-LCRAA-LCR7067.770.7OpenRouter model benchmarks
GDPval-AAGDPval-AA6854.655.1OpenRouter model benchmarks
CritPtCritPt29.120.916.9OpenRouter model benchmarks
SciCodeSciCode55.753.553.6OpenRouter model benchmarks
AA-Omniscience AccuracyAA-Omniscience Accuracy54.246.638.3OpenRouter model benchmarks
AA-Omniscience Non-Hallucination RateAA-Omniscience Non-Hallucination Rate49.964.162.7OpenRouter model benchmarks
IFBenchIFBench62.2OpenRouter model benchmarks
τ²-Bench Telecomτ²-Bench Telecom94.4OpenRouter model benchmarks
Terminal-Bench HardTerminal-Bench Hard58.3OpenRouter model benchmarks