Cohere model benchmarks.

Compare Cohere models across answer quality, coding, complex tasks, speed, price, and context size.

Cohere

Quality

Highest overall scores

  1. 1Cohere: Command A22.8
  2. 2Cohere: Command R+ (08-2024)0
  3. 3Cohere: Command R7B (12-2024)0

Response speed

Fastest output in this sample

  1. 1Cohere: Command A27 t/s
  2. 2Cohere: Command R7B (12-2024)20 t/s
  3. 3Cohere: Command R+ (08-2024)8 t/s

Input price

Lowest price per 1M tokens

  1. 1Cohere: Command R7B (12-2024)$0.04/1M
  2. 2Cohere: Command A$2.50/1M
  3. 3Cohere: Command R+ (08-2024)$2.50/1M

Highlights

Quality

Overall capability · Higher is better

Speed

Output tokens per second · Higher is better

Input price

USD per 1M tokens · Lower is better

Quality benchmarks

Compare Cohere model capability.

Performance and price

Compare practical tradeoffs.

All Cohere models.

3 models in the current catalogue.

Cohere technical benchmarks.

Compare named evaluation records available for Cohere models. Missing results remain blank rather than being estimated.

BenchmarkAreaCommand ACommand R+ (08-2024)Command R7B (12-2024)Source
Intelligence IndexoverallQuality22.8——OpenRouter / artificial-analysis
Coding Indexcoding27.8——OpenRouter / artificial-analysis
Agentic Indexagentic9.2——OpenRouter / artificial-analysis
GPQAGPQA Diamond76.1——OpenRouter model benchmarks
Humanity's Last ExamHLE12——OpenRouter model benchmarks
IFBenchIFBench73.9——OpenRouter model benchmarks
τ²-Bench Telecomτ²-Bench Telecom80.7——OpenRouter model benchmarks
AA-LCRAA-LCR48.7——OpenRouter model benchmarks
GDPval-AAGDPval-AA10.8——OpenRouter model benchmarks
CritPtCritPt0.3——OpenRouter model benchmarks
SciCodeSciCode37.8——OpenRouter model benchmarks
Terminal-Bench HardTerminal-Bench Hard25——OpenRouter model benchmarks
AA-Omniscience AccuracyAA-Omniscience Accuracy8.9——OpenRouter model benchmarks
AA-Omniscience Non-Hallucination RateAA-Omniscience Non-Hallucination Rate85.8——OpenRouter model benchmarks