AI model benchmarks.

Compare broadly useful AI models across intelligence, coding, tool use, response speed, price, context size, and availability. Every section explains what is measured and whether a higher or lower value is better.

Intelligence

Strongest performance across published evaluations

  1. 1Claude Opus 563.1
  2. 2xAI: Grok 4.660.9

Response speed

Fastest output in this sample

  1. 1MiniMax: MiniMax M2.7279 t/s
  2. 2Google: Gemini 3.5 Flash231.5 t/s

Cost per task

Lowest estimated cost for a typical task

  1. 1IBM: Granite 4.0 Micro$0.0001
  2. 2Upstage: Solar Pro 4$0.0001

Context size

Most content read at once

  1. 1OpenAI: GPT-5.6 Sol1.1M
  2. 2OpenAI: GPT-5.6 Terra1.1M

Highlights

Intelligence

Published evaluation score · Higher is better

Output speed

Generated tokens per second · Higher is better

Cost per task

Estimated cost for a typical task · Lower is better

Intelligence

How capable is each model?

Intelligence combines published evaluations into a comparable score. Use the tabs to focus on general capability, writing code, or completing tasks with tools. Higher is better.

Named evaluations

Explore technical benchmarks.

Select an evaluation such as GPQA, MMLU-Pro, Humanity's Last Exam, SWE-bench, Terminal-Bench, or AIME. Only models with a compatible published result are ranked.

RankModelCompanyScoreSource
1xAI: Grok 4.6xAI94.9OpenRouter model benchmarks
2OpenAI: GPT-5.6 SolOpenAI94.1OpenRouter model benchmarks
3Google: Gemini 3.1 Pro PreviewGoogle94.1OpenRouter model benchmarks
4Claude Opus 5Anthropic93.2OpenRouter model benchmarks
5SpaceXAI: Grok 4.5xAI93.1OpenRouter model benchmarks
6MiniMax: MiniMax M3MiniMax92.9OpenRouter model benchmarks
7Google: Gemini 3.6 FlashGoogle92.8OpenRouter model benchmarks
8Qwen: Qwen3.8 MaxAlibaba92.7OpenRouter model benchmarks
9OpenAI: GPT-5.6 TerraOpenAI92.5OpenRouter model benchmarks
10Google: Gemini 3.5 FlashGoogle92.2OpenRouter model benchmarks
11Anthropic: Claude Opus 4.8Anthropic92OpenRouter model benchmarks
12Anthropic: Claude Sonnet 5Anthropic91.1OpenRouter model benchmarks
13OpenAI: GPT-5.6 LunaOpenAI91.1OpenRouter model benchmarks
14DeepSeek: DeepSeek V4 Flash 0731DeepSeek90.8OpenRouter model benchmarks
15Qwen: Qwen3.7 PlusAlibaba90OpenRouter model benchmarks
16Meta: Muse Spark 1.1Meta89.8OpenRouter model benchmarks
17Tencent: Hy3 previewTencent89.7OpenRouter model benchmarks
18MoonshotAI: Kimi K2.7 CodeMoonshot AI89.6OpenRouter model benchmarks
19Z.ai: GLM 5.2Z.ai89.5OpenRouter model benchmarks
20Thinking Machines: Inkling SmallThinking Machines89.5OpenRouter model benchmarks
21Upstage: Solar Pro 4Upstage89.1OpenRouter model benchmarks
22DeepSeek: DeepSeek V4 Pro 0813DeepSeek88.8OpenRouter model benchmarks
23MiniMax: MiniMax M2.7MiniMax87.4OpenRouter model benchmarks
24Thinking Machines: InklingThinking Machines87.2OpenRouter model benchmarks
25Z.ai: GLM 5.1Z.ai86.8OpenRouter model benchmarks
26NVIDIA: Nemotron 3 UltraNVIDIA86.7OpenRouter model benchmarks
27Xiaomi: MiMo-V2.5Xiaomi84.9OpenRouter model benchmarks
28MiniMax: MiniMax M2.5MiniMax84.8OpenRouter model benchmarks
29StepFun: Step 3.5 FlashStepFun83.1OpenRouter model benchmarks
30Amazon: Nova 2 LiteAmazon81.1OpenRouter model benchmarks
31StepFun: Step 3.7 FlashStepFun80.9OpenRouter model benchmarks
32Mistral: Mistral Small 4Mistral AI76.9OpenRouter model benchmarks
33Cohere: Command ACohere76.1OpenRouter model benchmarks
34NVIDIA: Nemotron 3 Nano 30B A3BNVIDIA75.7OpenRouter model benchmarks
35Z.ai: GLM 4.7 FlashZ.ai75.2OpenEvals/leaderboard-data
36Mistral: Mistral Medium 3.5Mistral AI74.8OpenRouter model benchmarks
37NVIDIA: Nemotron 3.5 LightningNVIDIA74.3OpenRouter model benchmarks
38Upstage: Solar Pro 3Upstage72.4OpenRouter model benchmarks
39DeepSeek: DeepSeek V4 Flash 0423DeepSeek71.6OpenRouter model benchmarks
40MoonshotAI: Kimi K2 ThinkingMoonshot AI71.3OpenRouter model benchmarks
41Mistral: Mistral Large 3 2512Mistral AI68OpenRouter model benchmarks
42Perplexity: Sonar ProPerplexity57.8OpenRouter model benchmarks
43Microsoft: Phi 4Microsoft57.5OpenRouter model benchmarks
44Reka Flash 3Reka AI52.9OpenRouter model benchmarks
45Meta: Llama 3.3 70B InstructMeta49.8OpenRouter model benchmarks
46Amazon: Nova Lite 1.0Amazon43.3OpenRouter model benchmarks
47IBM: Granite 4.1 8BIBM43.3OpenRouter model benchmarks
48AI21: Jamba Large 1.7AI21 Labs39OpenRouter model benchmarks
49IBM: Granite 4.0 MicroIBM33.6OpenRouter model benchmarks

Intelligence comparisons

Compare capability with practical tradeoffs.

Start with intelligence, then inspect the practical measurement that matters to you. Lower is better for cost and answer-start time; higher is better for speed and context.

Token use

How long can one answer be?

This shows the maximum response length allowed by each model, not the number of tokens a typical answer will actually use. Higher means the model can produce a longer single response.

Cost

What does using each model cost?

Cost per task estimates 1,000 input tokens and 500 output tokens. Input and output prices show the underlying API price per one million tokens. Lower is better.

Context window

How much information can the model read at once?

The context window is the combined amount of instructions, conversation, documents, and generated text the model can consider in one request. Higher is better for large documents and long conversations.

Output speed

How quickly does the answer appear?

Output speed measures how many tokens arrive each second after the model begins responding. Higher is better. These are observed OpenRouter endpoint results, not a guarantee for every provider or request.

Answer start time

How long before the model starts replying?

This is the delay before the first part of an answer arrives. Lower is better. A model can start quickly but still generate the rest of its answer slowly.

Total response time

How long for a complete typical answer?

This estimate combines answer-start time with the time required to generate 500 output tokens. Lower is better. Actual time changes with answer length, provider load, and reasoning settings.

Model size

How large are open-weight models?

Parameter count is one rough indicator of the memory and hardware needed to run a model locally. It is not an intelligence score, and providers do not publish it for every model.

Prompt options

What can developers control?

These are common API controls supported by models in this catalogue. Availability can vary by provider. The count shows how many listed models support each option.

Max Tokens

Supported by 76 of 76 listed models

Temperature

Supported by 65 of 76 listed models

Tools

Supported by 65 of 76 listed models

Tool Choice

Supported by 63 of 76 listed models

Include Reasoning

Supported by 61 of 76 listed models

Reasoning

Supported by 61 of 76 listed models

Response Format

Supported by 61 of 76 listed models

Top P

Supported by 61 of 76 listed models

Structured Outputs

Supported by 60 of 76 listed models

Stop

Supported by 56 of 76 listed models

Seed

Supported by 52 of 76 listed models

Frequency Penalty

Supported by 50 of 76 listed models