Every answer on this page was written by an AI model, named on the answer. Answers can be wrong, out of date or made up, even when they sound sure. Check anything that matters before you rely on it.
On planning and architecture, Opus 5.5 at max effort scored highest, 97 out of 100. Three more were within 5 points: Fable 5.1 at max effort (94), Sonnet 5 at low effort (94) and Fable 5 at max effort (94). The lowest was Sonnet 5 at max effort, 63.
Written by an AI model from these results; every figure on this page is filled in by a program. How it was made.
Published
Score on planning and architecture
Each model and setting, highest first. A score here is the mean share of checks passed over the four prompts in this kind of work.
Score on planning and architecture
Light bars are low effort. Dark bars are max effort, or the only setting.
Show the numbers.
Score out of 100 on planning and architecture, for each model and setting, highest first
Model
Score
Opus 5.5 · max
97
Fable 5.1 · max
94
Sonnet 5 · low
94
Fable 5 · max
94
Fable 5.1 · low
91
Sonnet 5.5 · max
88
Sonnet 5.5 · low
85
Fable 5 · low
77
Haiku 4.5
77
Opus 5.5 · low
76
Opus 5 · max
76
Opus 5 · low
65
Sonnet 5 · max
63
Each prompt
Every cell is the share of that prompt’s checks the answer passed.
Score for each prompt in planning and architecture
Full marks
80 to 99
50 to 79
1 to 49
No checks passed
† The answer ran out of room: it reached the model’s output limit, mostly spent thinking, before it was finished. It is scored on what it did write.
Show the numbers.
Score out of 100 on each prompt in planning and architecture, for each model and setting
A developer is building online booking for their sister's one-person dog grooming business. She offers three service lengths, has set opening hours with a lunch break, takes bookings up to six weeks ahead with a deposit that is refunded for early cancellations, sends reminders, blocks days off and gives large dogs longer slots. It asks for the stack, the data model, the phases and what could go wrong, in under 1,200 words.
Every check passed: by 5 of 13 answers.
Missed most often:The first phase is small and usable before payments exist, in 7 of 13 answers.
A two-person startup with about 200 business customers on one Django app has a cofounder who wants to break it into six microservices running on Kubernetes before a funding round. Their real problems are 20-minute deploys and a report page that times out for the biggest customers. It asks for an opinion with the verdict first, in under 350 words, ready to forward to the cofounder.
Every check passed: by 4 of 13 answers.
Missed most often:Suggests a middle path: clear modules inside the monolith, in 6 of 13 answers.
Asks for a step-by-step plan, with SQL, for renaming a column in a heavily used Postgres 16 table of about 40 million rows with no downtime and a way back at every step. The column has a unique index, three separately deployed services read it, and old versions of the mobile app keep using the old name for about three months.
Every check passed: by 1 of 13 answers, Sonnet 5 at low effort.
Missed most often:Drops the old column last, after a waiting period, in 11 of 13 answers.
Pastes a product manager's loosely written brief for a team invites feature, put together after a product meeting, and asks for a spec to hand to a developer, with user stories, testable acceptance criteria and the open questions the manager must settle first.
Every check passed: by 3 of 13 answers, Opus 5.5 at max effort, Fable 5.1 at max effort and Sonnet 5 at low effort.
Missed most often:Flags the free plan limit of 3 seats versus 5 people, in 6 of 13 answers.
What max effort changed
Of the six models tested at both low and max effort, five scored higher here at max effort, one lower and no the same. The widest change was Sonnet 5, from 94 at low effort to 63 at max.
Score on planning and architecture at low and max effort, for each model with both
Each prompt asks for a plan rather than code: a booking app for a dog groomer, an answer on switching to microservices, a column rename with no downtime, and a spec drawn from a messy feature brief.
A program checks what can be counted, such as the parts asked for, the word limit and the steps a safe rename needs. A second model reads the rest against plain checks written before any model answered, such as whether the old column is dropped only after old app versions have stopped using it.
Four prompts and 34 checks in all: 17 run or counted by a program and 17 read by a second model. The full method.