StepFun model benchmarks.

Compare StepFun models across answer quality, coding, complex tasks, speed, price, and context size.

Quality

Highest overall scores

  1. 1StepFun: Step 3.7 Flash30.9
  2. 2StepFun: Step 3.5 Flash0

Response speed

Fastest output in this sample

  1. 1StepFun: Step 3.7 Flash159 t/s
  2. 2StepFun: Step 3.5 Flash54 t/s

Input price

Lowest price per 1M tokens

  1. 1StepFun: Step 3.5 Flash$0.10/1M
  2. 2StepFun: Step 3.7 Flash$0.20/1M

Highlights

Quality

Overall capability · Higher is better

Speed

Output tokens per second · Higher is better

Input price

USD per 1M tokens · Lower is better

Quality benchmarks

Compare StepFun model capability.

Performance and price

Compare practical tradeoffs.

All StepFun models.

2 models in the current catalogue.

ModelLinks
StepFun: Step 3.7 Flash30.9159 t/s262.1K$0.20$1.15
StepFun: Step 3.5 Flash054 t/s262.1K$0.10$0.30

StepFun technical benchmarks.

Compare named evaluation records available for StepFun models. Missing results remain blank rather than being estimated.

BenchmarkAreaStep 3.7 FlashStep 3.5 FlashSource
Intelligence IndexoverallQuality30.9—OpenRouter / artificial-analysis
Coding Indexcoding39.6—OpenRouter / artificial-analysis
Agentic Indexagentic21.7—OpenRouter / artificial-analysis
GPQAGPQA Diamond80.983.1OpenRouter model benchmarks
Humanity's Last ExamHLE21.421.1OpenRouter model benchmarks
IFBenchIFBench67.364.6OpenRouter model benchmarks
τ²-Bench Telecomτ²-Bench Telecom98.594.4OpenRouter model benchmarks
AA-LCRAA-LCR69.748OpenRouter model benchmarks
GDPval-AAGDPval-AA25.9—OpenRouter model benchmarks
CritPtCritPt2.32.5OpenRouter model benchmarks
SciCodeSciCode4040.4OpenRouter model benchmarks
Terminal-Bench HardTerminal-Bench Hard35.627.3OpenRouter model benchmarks
AA-Omniscience AccuracyAA-Omniscience Accuracy25.823.6OpenRouter model benchmarks
AA-Omniscience Non-Hallucination RateAA-Omniscience Non-Hallucination Rate1514.3OpenRouter model benchmarks
MMLU-ProKnowledge and reasoning—84.4OpenEvals/leaderboard-data
SWE-bench VerifiedCoding—74.4OpenEvals/leaderboard-data
Terminal-BenchAgentic coding—51OpenEvals/leaderboard-data
AIME 2026Math—96.67OpenEvals/leaderboard-data