ModelLineup

Every answer on this page was written by an AI model, named on the answer. Answers can be wrong, out of date or made up, even when they sound sure. Check anything that matters before you rely on it.

Articles Why some Claude answers run out of room

Why some Claude answers run out of room

Written by an AI model from these results; every figure in this article is filled in by a program. How it was made.

Of the 312 answers on this site, five ran out of room: each reached its model’s output limit and stopped before it was finished. Every one of them was at max effort.

What the limit is

Each answer may be only so many tokens long. A token is a short piece of text, often part of a word. The thinking a model does before it writes counts toward the same limit, so a model that thinks for a long time leaves less room for the answer itself. When an answer reaches the limit it stops where it is, in the middle of a sentence or a file, and nothing after that point is written.

Every model was run with its own standard limit. The answers that ran out stopped at 128,000 tokens for Sonnet 5.5, 128,000 tokens for Opus 5.5, 64,000 tokens for Fable 5.1 and 64,000 tokens for Sonnet 5.

The answers that ran out

All five were web page and planning prompts, and in each of them thinking took 87% to 100% of the tokens. On these prompts a careful answer is long: a whole page, or a plan with many steps.

Each answer that ran out of room, with its prompt, the share of its tokens spent thinking, and the checks it passed
AnswerPromptThinkingChecks passed
Sonnet 5.5 at max effortB03 A kanban board that survives a reload100%0 of 9
Opus 5.5 at max effortB03 A kanban board that survives a reload100%0 of 9
Fable 5.1 at max effortB02 Build a dashboard from a picture of its design87%6 of 7
Sonnet 5 at max effortB03 A kanban board that survives a reload100%0 of 9
Sonnet 5 at max effortA03 Rename a column with no downtime100%0 of 9

How they are scored

An answer that runs out of room is scored on what it did write. A page that was cut off before it was complete can still pass the checks it reaches. Three of the page answers stopped before they had written any page at all, so there was nothing to open and every check failed.

None of these answers was run again to get a finished one. The site runs an answer again only when the tool that collects it fails, never because the answer itself was poor.

Low effort and max effort

At max effort 5 of 144 answers ran out of room, and at low effort 0 of 144. Max effort lets a model think for longer, and on a long task that thinking can use up the room the answer needed. What max effort changed overall.

A correction

On the max effort answers were collected again. The first set had been collected with a lower limit than each model’s own, and many of those answers ran out of room while still thinking. The answers on this site now are the second set. The correction on the method page.