Every answer on this page was written by an AI model, named on the answer. Answers can be wrong, out of date or made up, even when they sound sure. Check anything that matters before you rely on it.
Articles/Does max effort help? What 144 pairs show
Does max effort help? What 144 pairs show
Published
Written by an AI model from these results; every figure in this article is filled in by a program. How it was made.
Six models answered each of the 24 prompts twice: once at low effort, and once at max effort, where a model may think for much longer before it writes. That gives 144 pairs of answers to the same question from the same model. Max effort scored higher on 36 of them, lower on 15, and the same on 93.
The overall scores
For five of the six models (Sonnet 5.5, Opus 5.5, Fable 5.1, Opus 5 and Fable 5), the overall score moved by less than 5 points, which the site counts as level. Sonnet 5 scored lower at max effort by 5 points or more.
Low effort against max effort
Ring: low effort. Dot: max effort.
Show the numbers.
Score out of 100 at low effort (ring) and max effort (dot), for each model with both
Model
Low effort
Max effort
Fable 5.1
95
97
Fable 5
91
94
Opus 5.5
90
93
Sonnet 5.5
93
92
Opus 5
87
91
Sonnet 5
88
83
Where it helped, and where it did not
By kind of work, averaged over the six models, the score at max effort against low effort was:
Over all 24 answers, max effort cost between 3.3 and 16.5 times as much as low effort, depending on the model, at API prices, and took between 3.8 and 18.9 times as long. Thinking is most of the extra: the six models wrote 3.3M thinking tokens at max effort, against 84k at low, out of 3.6M and 340k output tokens in all.
Score at low and max effort, the change in points, and how many times the cost and the time max effort took, for each model
Longer thinking uses up more of an answer’s output limit. At max effort 5 answers ran out of room before they were finished, against 0 at low effort. Why some answers run out of room.
What this does not show
Each model answered each prompt once at each setting, so a single answer can move a prompt’s score, and a gap of less than 5 points overall is treated as level. The prompts are everyday and working tasks, not puzzles built to need long reasoning. The effort charts for each model and the method have the detail.