ModelLineup

Every answer on this page was written by an AI model, named on the answer. Answers can be wrong, out of date or made up, even when they sound sure. Check anything that matters before you rely on it.

Articles Does max effort help? What 144 pairs show

Does max effort help? What 144 pairs show

Written by an AI model from these results; every figure in this article is filled in by a program. How it was made.

Six models answered each of the 24 prompts twice: once at low effort, and once at max effort, where a model may think for much longer before it writes. That gives 144 pairs of answers to the same question from the same model. Max effort scored higher on 36 of them, lower on 15, and the same on 93.

The overall scores

For five of the six models (Sonnet 5.5, Opus 5.5, Fable 5.1, Opus 5 and Fable 5), the overall score moved by less than 5 points, which the site counts as level. Sonnet 5 scored lower at max effort by 5 points or more.

Low effort against max effort
Low effort against max effort0255075100Fable 5.1: 95 at low effort, 97 at max effortFable 5.197Fable 5: 91 at low effort, 94 at max effortFable 594Opus 5.5: 90 at low effort, 93 at max effortOpus 5.593Sonnet 5.5: 93 at low effort, 92 at max effortSonnet 5.592Opus 5: 87 at low effort, 91 at max effortOpus 591Sonnet 5: 88 at low effort, 83 at max effortSonnet 583Low effort against max effort0255075100Fable 5.1: 95 at low effort, 97 at max effortFable 5.197Fable 5: 91 at low effort, 94 at max effortFable 594Opus 5.5: 90 at low effort, 93 at max effortOpus 5.593Sonnet 5.5: 93 at low effort, 92 at max effortSonnet 5.592Opus 5: 87 at low effort, 91 at max effortOpus 591Sonnet 5: 88 at low effort, 83 at max effortSonnet 583

Ring: low effort. Dot: max effort.

Show the numbers.
Score out of 100 at low effort (ring) and max effort (dot), for each model with both
ModelLow effortMax effort
Fable 5.19597
Fable 59194
Opus 5.59093
Sonnet 5.59392
Opus 58791
Sonnet 58883

Where it helped, and where it did not

By kind of work, averaged over the six models, the score at max effort against low effort was:

A prompt’s score is the share of its checks passed, out of 100. Averaged the same way, the prompts that moved most:

What it cost

Over all 24 answers, max effort cost between 3.3 and 16.5 times as much as low effort, depending on the model, at API prices, and took between 3.8 and 18.9 times as long. Thinking is most of the extra: the six models wrote 3.3M thinking tokens at max effort, against 84k at low, out of 3.6M and 340k output tokens in all.

Score at low and max effort, the change in points, and how many times the cost and the time max effort took, for each model
ModelLowMaxChangeCostTime
Fable 5.19597+27.4 times8.0 times
Fable 59194+33.9 times4.6 times
Opus 5.59093+316.5 times18.9 times
Sonnet 5.59392-115.5 times18.0 times
Opus 58791+43.3 times3.8 times
Sonnet 58883-58.4 times10.8 times

Running out of room

Longer thinking uses up more of an answer’s output limit. At max effort 5 answers ran out of room before they were finished, against 0 at low effort. Why some answers run out of room.

What this does not show

Each model answered each prompt once at each setting, so a single answer can move a prompt’s score, and a gap of less than 5 points overall is treated as level. The prompts are everyday and working tasks, not puzzles built to need long reasoning. The effort charts for each model and the method have the detail.