Moonshot AI model benchmarks.

Compare Moonshot AI models across answer quality, coding, complex tasks, speed, price, and context size.

Quality

Highest overall scores

  1. 1MoonshotAI: Kimi K359.7
  2. 2MoonshotAI: Kimi K2.7 Code43
  3. 3MoonshotAI: Kimi K2 Thinking0

Response speed

Fastest output in this sample

  1. 1MoonshotAI: Kimi K2 Thinking148.5 t/s
  2. 2MoonshotAI: Kimi K2.7 Code118 t/s
  3. 3MoonshotAI: Kimi K365 t/s

Input price

Lowest price per 1M tokens

  1. 1MoonshotAI: Kimi K2 Thinking$0.60/1M
  2. 2MoonshotAI: Kimi K2.7 Code$0.67/1M
  3. 3MoonshotAI: Kimi K3$3.00/1M

Highlights

Quality

Overall capability · Higher is better

Speed

Output tokens per second · Higher is better

Input price

USD per 1M tokens · Lower is better

Quality benchmarks

Compare Moonshot AI model capability.

Performance and price

Compare practical tradeoffs.

All Moonshot AI models.

3 models in the current catalogue.

Moonshot AI technical benchmarks.

Compare named evaluation records available for Moonshot AI models. Missing results remain blank rather than being estimated.

BenchmarkAreaKimi K3Kimi K2.7 CodeKimi K2 ThinkingSource
Intelligence IndexoverallQuality59.743—OpenRouter / artificial-analysis
Coding Indexcoding76.260.821OpenRouter / artificial-analysis
Agentic Indexagentic54.330.3—OpenRouter / artificial-analysis
GPQAGPQA Diamond—89.671.3OpenRouter model benchmarks
Humanity's Last ExamHLE—3510.1OpenRouter model benchmarks
IFBenchIFBench—63.162.8OpenRouter model benchmarks
τ²-Bench Telecomτ²-Bench Telecom—90.125.4OpenRouter model benchmarks
AA-LCRAA-LCR—7555.3OpenRouter model benchmarks
GDPval-AAGDPval-AA—34.50OpenRouter model benchmarks
CritPtCritPt—100OpenRouter model benchmarks
SciCodeSciCode—47.533OpenRouter model benchmarks
Terminal-Bench HardTerminal-Bench Hard—44.76.8OpenRouter model benchmarks
AA-Omniscience AccuracyAA-Omniscience Accuracy—39.618OpenRouter model benchmarks
AA-Omniscience Non-Hallucination RateAA-Omniscience Non-Hallucination Rate—17.628.8OpenRouter model benchmarks
MMLU-ProKnowledge and reasoning——84.6OpenEvals/leaderboard-data
SWE-bench VerifiedCoding——71.3OpenEvals/leaderboard-data
Terminal-BenchAgentic coding——35.7OpenEvals/leaderboard-data