01
Remove models that do not fit
For example, a model must support tools before we can recommend it for taking actions on your behalf.
You choose the main thing you want to use AI for. We compare models using tests that match that work, remove models that do not have enough reliable data, and rank the rest. The model at the top has the best score for that use case. Popularity and company name do not affect the result.
01
For example, a model must support tools before we can recommend it for taking actions on your behalf.
02
Each use case has its own set of tests. Coding tests matter for writing code. Factual accuracy matters more for everyday help.
03
We rank models by the final score. When two models are almost tied, we check a relevant secondary test, the strength of the evidence, intelligence, speed, and cost.
Brainpower is a general score from 0 to 100. On the subscription table, it describes the strongest model included with that plan. A higher score means the model performed better across several kinds of difficult work.
Brainpower does not measure the plan's usage limit, price, features, or speed. It also does not pick every recommendation. A model with the highest Brainpower score may still lose for coding or taking actions if another model performs better on the tests for that job.
We show Not measured when the strongest included model is missing one of the benchmark groups needed for the calculation.
| Part of the score | Weight |
|---|---|
| Agents and tool use | 34% |
| Coding | 24% |
| Science and maths | 24% |
| General knowledge | 18% |
The percentages in each row add up to 100% because each row is a complete scoring recipe. They do not compare one use case with another. If you choose Writing code, we use the Writing code row. If you choose Everyday tasks, we use the Everyday tasks row.
A higher percentage means that test has more influence on the score for that use case. It does not mean that the model scored that percentage.
| Use case | How the score is built | A model must have |
|---|---|---|
| Everyday tasks | General intelligence 35% · GDPval-AA 20% · AA-Omniscience Accuracy 25% · IFBench 15% · AA-Omniscience Non-Hallucination Rate 5% | The required benchmark results shown in this row |
| Taking actions for you | Agentic Index 70% · Terminal-Bench Hard 10% · τ²-Bench Telecom 10% · IFBench 5% · General intelligence 5% | Tool support |
| Writing code | Coding Index 40% · Terminal-Bench Hard 25% · SciCode 15% · IFBench 10% · General intelligence 10% | The required benchmark results shown in this row |
| Math and logic | GPQA 45% · CritPt 25% · Humanity's Last Exam 15% · General intelligence 15% | The required benchmark results shown in this row |
Image quality comes from blind comparisons in the Artificial Analysis Text-to-Image Arena. We do not treat OpenRouter popularity as proof of quality.
We turn each model's leaderboard position into a score from 0 to 100. The number of comparisons and the margin of error tell us how certain that result is. Price matters only when quality scores are very close.
A result belongs to the exact model and quality setting that was tested. We do not copy a score from one setting to another.
We look at two jobs: making a video from written instructions and animating an existing image. Each job counts equally.
We score the two Artificial Analysis leaderboards separately and then combine the results. If a model was tested for only one job, we lower its confidence and subtract 15 points.
The combined score is calculated by 100x AI Models. We also show the original scores so you can see how the model performed in each job.
We calculate one score for every eligible model using the percentages in the matching row above. A coding score, for example, comes mostly from coding and terminal tests. An everyday score gives more weight to general intelligence and factual accuracy.
We then rank the models from highest to lowest. This is why a model can lead for Writing code but rank lower for Everyday tasks. The two choices use different tests and different weights.
A model needs every test marked as required and at least 70% of the evidence expected for that use case. Missing evidence lowers confidence. We do not replace a missing result with zero or an invented average.
The final score is the benchmark score minus an uncertainty penalty. Less complete or less reliable evidence can reduce the score by up to 15 points.
For some use cases, scores within one point are close enough to need another check. We then use a benchmark that is especially relevant to that use case. This stops a tiny difference in the main score from deciding the result on its own.
We do not guess a missing score from downloads, popularity, model size, marketing claims, or the company that made the model. If an important result is missing, that model does not appear in the recommendation for that use case. It still appears in the full benchmark catalogue.
Prices, context limits, features, and measured response speed come from our reviewed catalogue. Each benchmark keeps its source and version. We review catalogue updates before they appear on the site.
We review image and video leaderboard data by hand because the source does not provide a stable data feed. A media model needs an exact leaderboard match before we use that result in a recommendation.
The ranking answers the choices you made on this site. A different prompt, provider setting, or task may produce a different result.