DeepSeek model benchmarks.

Compare DeepSeek models across answer quality, coding, complex tasks, speed, price, and context size.

Quality

Highest overall scores

  1. 1DeepSeek: DeepSeek V4 Flash 073151.8
  2. 2DeepSeek: DeepSeek V4 Pro 081345.3
  3. 3DeepSeek: DeepSeek V4 Flash 04230

Response speed

Fastest output in this sample

  1. 1DeepSeek: DeepSeek V4 Flash 0731153 t/s
  2. 2DeepSeek: DeepSeek V4 Flash 042381 t/s
  3. 3DeepSeek: DeepSeek V4 Pro 081348 t/s

Input price

Lowest price per 1M tokens

  1. 1DeepSeek: DeepSeek V4 Flash 0731$0.08/1M
  2. 2DeepSeek: DeepSeek V4 Flash 0423$0.14/1M
  3. 3DeepSeek: DeepSeek V4 Pro 0813$0.43/1M

Highlights

Quality

Overall capability · Higher is better

Speed

Output tokens per second · Higher is better

Input price

USD per 1M tokens · Lower is better

Quality benchmarks

Compare DeepSeek model capability.

Performance and price

Compare practical tradeoffs.

All DeepSeek models.

3 models in the current catalogue.

DeepSeek technical benchmarks.

Compare named evaluation records available for DeepSeek models. Missing results remain blank rather than being estimated.

BenchmarkAreaDeepSeek V4 Flash 0731DeepSeek V4 Pro 0813DeepSeek V4 Flash 0423Source
Intelligence IndexoverallQuality51.845.3—OpenRouter / artificial-analysis
Coding Indexcoding69.159.4—OpenRouter / artificial-analysis
Agentic Indexagentic48.437.8—OpenRouter / artificial-analysis
GPQAGPQA Diamond90.888.871.6OpenRouter model benchmarks
Humanity's Last ExamHLE38.637.57.8OpenRouter model benchmarks
AA-LCRAA-LCR74.37037.3OpenRouter model benchmarks
GDPval-AAGDPval-AA52.940.3—OpenRouter model benchmarks
CritPtCritPt16.612.90.3OpenRouter model benchmarks
SciCodeSciCode49.95037.3OpenRouter model benchmarks
AA-Omniscience AccuracyAA-Omniscience Accuracy40.442.926OpenRouter model benchmarks
AA-Omniscience Non-Hallucination RateAA-Omniscience Non-Hallucination Rate8.35.95.5OpenRouter model benchmarks
IFBenchIFBench—76.547.2OpenRouter model benchmarks
τ²-Bench Telecomτ²-Bench Telecom—96.294.4OpenRouter model benchmarks
Terminal-Bench HardTerminal-Bench Hard—46.234.1OpenRouter model benchmarks