ModelLineup

Every answer on this page was written by an AI model, named on the answer. Answers can be wrong, out of date or made up, even when they sound sure. Check anything that matters before you rely on it.

Max effort results are being collected again. We capped every answer at 32,000 output tokens, below the output limit of six of the seven models, so many max effort answers ran out of room while still thinking and scored zero. Low effort results and Haiku 4.5 are not affected. Ranks are hidden until the new answers are in.

The same 24 questions, put to each model.

Real answers from Claude models to one fixed list of everyday and working questions, shown in full next to plain checks, time, tokens and cost.

Sonnet 5.5

24 prompts, every check counted

Low effort

93

  • Floor45/49, 92% passed
  • Middle53/58, 91% passed
  • Top62/65, 95% passed

Median answer 8.8 s, 1.7¢.

Max effort

49

  • Floor29/49, 59% passed
  • Middle28/58, 48% passed
  • Top32/65, 49% passed

Median answer 3 min 19 s, 30¢.

Max passed 71 fewer checks, cost 9.4 times as much and took 11.3 times as long. Costs are at Anthropic’s API prices on .

Every answer is marked in three layers

Each prompt has a few plain checks, written before any model answered. A score is the share of checks an answer passes, so a longer answer earns nothing extra.

  • Floor29 of 49, 59% passed
  • Middle28 of 58, 48% passed
  • Top32 of 65, 49% passed

Sonnet 5.5, all 24 prompts

  1. Floor 49 checks

    What nearly every answer should get right: the page loads, the number is there, the question was answered.

  2. Middle 58 checks

    What a careful answer gets right: the edge case, the rule from the brief, the detail that is easy to skip.

  3. Top 65 checks

    What only a strong answer manages: the hard part, done properly, with nothing left over.

172 checks in all. 136 are counted or tested by a program, and 36 need reading, which a second model does, twice. The full method.

The models tested

Each model is tested at its lowest and highest effort where it has the setting. Open one to read every answer.

Released

Haiku 4.5

72/ 100

  • One setting, thinking off 72
  • Floor47/49, 96% passed
  • Middle45/58, 78% passed
  • Top32/65, 49% passed

Tested . Strongest in Planning and architecture.

Released

Fable 5

95/ 100

  • Low effort 91
  • Max effort 95
  • Floor48/49, 98% passed
  • Middle56/58, 97% passed
  • Top60/65, 92% passed

Tested . Strongest in Code.

Released

Sonnet 5

54/ 100

  • Low effort 88
  • Max effort 54
  • Floor29/49, 59% passed
  • Middle30/58, 52% passed
  • Top34/65, 52% passed

Tested . Strongest in Everyday.

Released

Opus 5

91/ 100

  • Low effort 87
  • Max effort 91
  • Floor45/49, 92% passed
  • Middle52/58, 90% passed
  • Top59/65, 91% passed

Tested . Strongest in Code.

Released

Fable 5.1

72/ 100

  • Low effort 95
  • Max effort 72
  • Floor39/49, 80% passed
  • Middle42/58, 72% passed
  • Top45/65, 69% passed

Tested . Strongest in Everyday.

Released

Opus 5.5

49/ 100

  • Low effort 90
  • Max effort 49
  • Floor28/49, 57% passed
  • Middle28/58, 48% passed
  • Top32/65, 49% passed

Tested . Strongest in Everyday.

Released

Sonnet 5.5

49/ 100

  • Low effort 93
  • Max effort 49
  • Floor29/49, 59% passed
  • Middle28/58, 48% passed
  • Top32/65, 49% passed

Tested . Strongest in Everyday.

How the answers were collected

The instruction, the checks, the judge and the prices used, in one place.

Read the method