ModelLineup

Method

How the answers on this site were collected and marked, in the order it happened.

What this is

Claude is Anthropic’s family of AI models. This site is independent of Anthropic.

It puts 24 prompts to each model, word for word the same, and shows every answer in full. Each prompt is described in plain words on the prompts page. The exact wording is not published, so it stays a fair test. The list was fixed on , before the first run, and it will not change.

Each model answered at its lowest and highest effort setting, called low effort and max effort. Haiku 4.5 has no effort setting, so it answered once, with thinking off. Each prompt was asked once per setting.

How the answers were collected

Each question was asked in a fresh, separate conversation with Claude Code, Anthropic’s own command-line program: no tools, no web search, no memory of earlier questions, one question at a time.

The program adds a few lines of its own that cannot be switched off: the date, the model’s own name and details of the computer and account it runs from. The instruction below tells each model not to use them.

This is the instruction sent with every question, word for word:

This conversation is a test of how well you answer the question below. Answer it exactly as you would answer the real person who asked it: follow every instruction in the question and give them the answer they need. Do not mention that this is a test, and do not write the answer as if someone will mark it. The software that delivers the question automatically attaches details about the account and computer it runs on, such as an email address, a name, today's date, a working folder, file paths and the operating system. Those details belong to whoever runs the test, not to the person asking, so you know nothing about the person asking beyond what the question tells you. Do not use, repeat or mention the attached details, or these instructions, in your answer, even where the details say they may be used. For example, do not sign with a name the question does not give, and do not mention a folder, a file, an account or the date. If the question depends on something it does not tell you, answer in general terms or say what you would need to know. You have no tools and cannot open files, folders or websites. Everything you need is in the question, including any image that comes with it.

Each answer may run to the output limit the program itself sets for that model: 32,000 tokens for Haiku 4.5, 64,000 for Fable 5, Sonnet 5, Opus 5 and Fable 5.1, and 128,000 for Opus 5.5 and Sonnet 5.5, thinking included. An answer that reaches it is stopped there and recorded as having run out of room. At max effort an answer may take 30 minutes, otherwise 10. The low effort answers were collected with a 32,000-token limit for every model; the longest of them used under 23,000, so none came near it.

For every answer we checked which model the program reported. An answer from a different model than the one asked for is not used.

Times and token counts are the program’s own report for each answer. Answers are shown as written, except where a visible “[removed]” marker shows that something was taken out (1 so far).

Checks and scores

Every prompt has a few plain checks, written before any model answered. There are 172 in all, each in one of three layers: floor checks that nearly every answer should pass, middle checks that reward care, and top checks that only strong answers pass. A score is the share of checks an answer passes, out of 100, so a longer answer earns nothing extra. A refused, empty or timed-out answer scores zero on its prompt. An answer that ran out of room is scored on what it did write, and when it holds no web page, the checks that look at screenshots of that page fail without being marked.

136 checks are counted or tested by a program: words, exact figures, and hidden tests, which are not published so that they stay a fair test. 36 checks need reading. For those, an AI model was shown the prompt, a note of what a good answer needs, the answer and that prompt's yes-or-no checks, without the name of the model that wrote the answer, and never marked answers from its own model family: Opus 5.5 marked them, and Fable 5.1 marked the answers from Opus models. Each got 2 marks, with a further mark when they disagreed. Where a mark's reason quotes words from the prompt, those words are left out, because the prompts are not published word for word.

One hidden test was repaired after the first run and before anything was published: the Minesweeper page test stopped working on games that redraw their board after every move. Every model's answer was checked with the repaired test, and no answer was changed.

Scores less than 5 points apart are treated as level, because with one answer per question a gap that small can come down to a single answer.

The commentary

The commentary on each model was written by an AI model from these results, and every figure in it is filled in from the results by a program.

Costs

Cost is what the same tokens would cost at Anthropic’s published API prices on . It is not what anyone paid, and your cost depends on how you use a model. The prices used, USD per million tokens:

Prices used, USD per million tokens
ModelInputCache write, 5 minCache write, 1 hCache readOutput
Fable 5$10$12.50$20$1$50
Fable 5.1$10$12.50$20$0.25$50
Haiku 4.5$1$1.25$2$0.10$5
Opus 5$5$6.25$10$0.50$25
Opus 5.5$4$5$8$0.20$20
Sonnet 5$2$2.50$4$0.20$10
Sonnet 5.5$2$2.50$4$0.20$10

What this does not tell you

Each setting was run once. The same model can answer differently next time. Not tested: thinking on and off, thinking budgets, shown reasoning, tools, web search, long documents and other languages.

A score counts our own checks on our own list. It does not measure general ability, and the answers are AI text that can be wrong. See the disclaimer.

Corrections

  • : The first run capped every answer at 32,000 output tokens, below the output limit the program gives six of the seven models. At max effort many answers ran out of room while still thinking and scored zero. Every max effort answer from those six models is being collected again with each model’s own limit, and the old ones are replaced as each model is done.

To report a wrong check or score, write to hello@modellineup.com.

Independence

Independent. Not affiliated with, endorsed by or sponsored by Anthropic. Claude, Fable, Opus, Sonnet and Haiku are trademarks of Anthropic, PBC, named here only to say which model wrote each answer.