How we choose a model.

You choose the main thing you want to use AI for. We compare models using tests that match that work, remove models that do not have enough reliable data, and rank the rest. The model at the top has the best score for that use case. Popularity and company name do not affect the result.

01

Remove models that do not fit

For example, a model must support tools before we can recommend it for taking actions on your behalf.

02

Score the work you selected

Each use case has its own set of tests. Coding tests matter for writing code. Factual accuracy matters more for everyday help.

03

Put the best match first

We rank models by the final score. When two models are almost tied, we check a relevant secondary test, the strength of the evidence, intelligence, speed, and cost.

What Brainpower means.

Brainpower is a general score from 0 to 100. On the subscription table, it describes the strongest model included with that plan. A higher score means the model performed better across several kinds of difficult work.

Brainpower does not measure the plan's usage limit, price, features, or speed. It also does not pick every recommendation. A model with the highest Brainpower score may still lose for coding or taking actions if another model performs better on the tests for that job.

We show Not measured when the strongest included model is missing one of the benchmark groups needed for the calculation.

Part of the scoreWeight
Agents and tool use34%
Coding24%
Science and maths24%
General knowledge18%

What we measure for each use case.

The percentages in each row add up to 100% because each row is a complete scoring recipe. They do not compare one use case with another. If you choose Writing code, we use the Writing code row. If you choose Everyday tasks, we use the Everyday tasks row.

A higher percentage means that test has more influence on the score for that use case. It does not mean that the model scored that percentage.

Use caseHow the score is builtA model must have
Everyday tasksGeneral intelligence 35% · GDPval-AA 20% · AA-Omniscience Accuracy 25% · IFBench 15% · AA-Omniscience Non-Hallucination Rate 5%The required benchmark results shown in this row
Taking actions for youAgentic Index 70% · Terminal-Bench Hard 10% · τ²-Bench Telecom 10% · IFBench 5% · General intelligence 5%Tool support
Writing codeCoding Index 40% · Terminal-Bench Hard 25% · SciCode 15% · IFBench 10% · General intelligence 10%The required benchmark results shown in this row
Math and logicGPQA 45% · CritPt 25% · Humanity's Last Exam 15% · General intelligence 15%The required benchmark results shown in this row

How image and video models are ranked.

Image generation.

Image quality comes from blind comparisons in the Artificial Analysis Text-to-Image Arena. We do not treat OpenRouter popularity as proof of quality.

We turn each model's leaderboard position into a score from 0 to 100. The number of comparisons and the margin of error tell us how certain that result is. Price matters only when quality scores are very close.

A result belongs to the exact model and quality setting that was tested. We do not copy a score from one setting to another.

Video generation.

We look at two jobs: making a video from written instructions and animating an existing image. Each job counts equally.

We score the two Artificial Analysis leaderboards separately and then combine the results. If a model was tested for only one job, we lower its confidence and subtract 15 points.

The combined score is calculated by 100x AI Models. We also show the original scores so you can see how the model performed in each job.

How your selected use case is scored.

We calculate one score for every eligible model using the percentages in the matching row above. A coding score, for example, comes mostly from coding and terminal tests. An everyday score gives more weight to general intelligence and factual accuracy.

We then rank the models from highest to lowest. This is why a model can lead for Writing code but rank lower for Everyday tasks. The two choices use different tests and different weights.

A model needs every test marked as required and at least 70% of the evidence expected for that use case. Missing evidence lowers confidence. We do not replace a missing result with zero or an invented average.

The final score is the benchmark score minus an uncertainty penalty. Less complete or less reliable evidence can reduce the score by up to 15 points.

For some use cases, scores within one point are close enough to need another check. We then use a benchmark that is especially relevant to that use case. This stops a tiny difference in the main score from deciding the result on its own.

What happens when data is missing.

We do not guess a missing score from downloads, popularity, model size, marketing claims, or the company that made the model. If an important result is missing, that model does not appear in the recommendation for that use case. It still appears in the full benchmark catalogue.

Prices, context limits, features, and measured response speed come from our reviewed catalogue. Each benchmark keeps its source and version. We review catalogue updates before they appear on the site.

We review image and video leaderboard data by hand because the source does not provide a stable data feed. A media model needs an exact leaderboard match before we use that result in a recommendation.

The ranking answers the choices you made on this site. A different prompt, provider setting, or task may produce a different result.