72/ 100
- One setting, thinking off 72
- Floor47/49, 96% passed
- Middle45/58, 78% passed
- Top32/65, 49% passed
Tested . Strongest in Planning and architecture.
Every answer on this page was written by an AI model, named on the answer. Answers can be wrong, out of date or made up, even when they sound sure. Check anything that matters before you rely on it.
Max effort results are being collected again. We capped every answer at 32,000 output tokens, below the output limit of six of the seven models, so many max effort answers ran out of room while still thinking and scored zero. Low effort results and Haiku 4.5 are not affected. Ranks are hidden until the new answers are in.
Real answers from Claude models to one fixed list of everyday and working questions, shown in full next to plain checks, time, tokens and cost.
Sonnet 5.5
24 prompts, every check counted
Low effort
93
Median answer 8.8 s, 1.7¢.
Max effort
49
Median answer 3 min 19 s, 30¢.
Max passed 71 fewer checks, cost 9.4 times as much and took 11.3 times as long. Costs are at Anthropic’s API prices on .
The list never changes. Every model gets the same 24 prompts, word for word, so the answers can be laid next to each other. Each prompt is described in plain words on the prompts page.
Each prompt has a few plain checks, written before any model answered. A score is the share of checks an answer passes, so a longer answer earns nothing extra.
What nearly every answer should get right: the page loads, the number is there, the question was answered.
What a careful answer gets right: the edge case, the rule from the brief, the detail that is easy to skip.
What only a strong answer manages: the hard part, done properly, with nothing left over.
172 checks in all. 136 are counted or tested by a program, and 36 need reading, which a second model does, twice. The full method.
Each model is tested at its lowest and highest effort where it has the setting. Open one to read every answer.
72/ 100
Tested . Strongest in Planning and architecture.
95/ 100
Tested . Strongest in Code.
54/ 100
Tested . Strongest in Everyday.
91/ 100
Tested . Strongest in Code.
72/ 100
Tested . Strongest in Everyday.
49/ 100
Tested . Strongest in Everyday.
49/ 100
Tested . Strongest in Everyday.
The instruction, the checks, the judge and the prices used, in one place.