Every answer on this page was written by an AI model, named on the answer. Answers can be wrong, out of date or made up, even when they sound sure. Check anything that matters before you rely on it.
Compare two models
Pick two of the seven models. Their scores, each kind of work, time and cost sit side by side, with the prompts where one passed checks the other missed.
Pick two different models.
Sonnet 5.5 and Opus 5.5
At max effort, Sonnet 5.5 scored 92 and Opus 5.5 93, 1.2 points apart, under the 5 points the site counts as level. At low effort, Sonnet 5.5 scored 93 and Opus 5.5 90, 2.4 points apart, under the 5 points the site counts as level.
| Sonnet 5.5 | Opus 5.5 | |||
|---|---|---|---|---|
| Low | Max | Low | Max | |
| Score | 93 | 92 | 90 | 93 |
| Web pages | 98 | 80 | 90 | 80 |
| Code | 98 | 100 | 98 | 100 |
| Planning | 85 | 88 | 76 | 97 |
| Everyday | 89 | 100 | 96 | 96 |
| Checks passed | 160 of 172 | 159 of 172 | 157 of 172 | 160 of 172 |
| Median time | 8.8 s | 2 min 17 s | 16 s | 5 min 14 s |
| Median cost | 1.7¢ | 20¢ | 4.1¢ | 76¢ |
Scored higher by 5 points or more: Opus 5.5 on planning and architecture, 97 to 88. Within 5 points on building web pages, code and everyday.
Sonnet 5.5’s median max effort answer took 2 min 17 s and cost 20¢ at API prices, and Opus 5.5’s median max effort answer took 5 min 14 s and cost 76¢. A median Opus 5.5 answer cost 3.7 times as much.
Of the 24 prompts at max effort, Sonnet 5.5 passed more checks on 2, Opus 5.5 on 2, and the two passed the same number on 20.
Where they differed
Prompts where the two passed a different number of checks at max effort, widest gap first.
A04 Turn a messy feature brief into a spec
Sonnet 5.5 6 of 8, Opus 5.5 8 of 8.
Passed by Opus 5.5, missed by Sonnet 5.5: Has user stories and acceptance criteria; Asks about an invitee already in another team and about pending invites on downgrade
R02 Did cycling in town really double?
Sonnet 5.5 6 of 6, Opus 5.5 5 of 6.
Passed by Sonnet 5.5, missed by Opus 5.5: Under 250 words
S02 A story of exactly 100 words
Sonnet 5.5 6 of 6, Opus 5.5 5 of 6.
Passed by Sonnet 5.5, missed by Opus 5.5: Uses the words 'kettle' and 'fog', and names the lighthouse or its keeper
A01 Plan a booking app for a dog groomer
Sonnet 5.5 8 of 9, Opus 5.5 9 of 9.
Passed by Opus 5.5, missed by Sonnet 5.5: The first phase is small and usable before payments exist
The full comparisonEvery answer from Sonnet 5.5Every answer from Opus 5.5
Sonnet 5.5 and Fable 5.1
At max effort, Sonnet 5.5 scored 92 and Fable 5.1 97, 4.8 points apart, under the 5 points the site counts as level. At low effort, Sonnet 5.5 scored 93 and Fable 5.1 95, 2.5 points apart, under the 5 points the site counts as level.
| Sonnet 5.5 | Fable 5.1 | |||
|---|---|---|---|---|
| Low | Max | Low | Max | |
| Score | 93 | 92 | 95 | 97 |
| Web pages | 98 | 80 | 98 | 95 |
| Code | 98 | 100 | 100 | 100 |
| Planning | 85 | 88 | 91 | 94 |
| Everyday | 89 | 100 | 92 | 98 |
| Checks passed | 160 of 172 | 159 of 172 | 164 of 172 | 167 of 172 |
| Median time | 8.8 s | 2 min 17 s | 20 s | 2 min 15 s |
| Median cost | 1.7¢ | 20¢ | 11¢ | 61¢ |
Scored higher by 5 points or more: Fable 5.1 on building web pages, 95 to 80; Fable 5.1 on planning and architecture, 94 to 88. Within 5 points on code and everyday.
Sonnet 5.5’s median max effort answer took 2 min 17 s and cost 20¢ at API prices, and Fable 5.1’s median max effort answer took 2 min 15 s and cost 61¢. A median Fable 5.1 answer cost 3.0 times as much.
Of the 24 prompts at max effort, Sonnet 5.5 passed more checks on 3, Fable 5.1 on 2, and the two passed the same number on 19.
Where they differed
Prompts where the two passed a different number of checks at max effort, widest gap first.
B03 A kanban board that survives a reload
Sonnet 5.5 0 of 9, Fable 5.1 9 of 9.
Passed by Fable 5.1, missed by Sonnet 5.5: The page loads with no errors and has the three named columns; Adding a card works with Enter and with the Add card button; Empty or space-only titles are ignored; The counts update, and Edit and Delete work; After a reload the board is exactly as it was; Dragging moves cards between columns and reorders them, and it survives a reload; A card can be moved with the arrow keys and keeps focus; Titles like <b>hi</b> & <img ...> show exactly as typed and run nothing; No sideways scroll on a 375 px phone, even with very long titles
A04 Turn a messy feature brief into a spec
Sonnet 5.5 6 of 8, Fable 5.1 8 of 8.
Passed by Fable 5.1, missed by Sonnet 5.5: Has user stories and acceptance criteria; Asks about an invitee already in another team and about pending invites on downgrade
R02 Did cycling in town really double?
Sonnet 5.5 6 of 6, Fable 5.1 5 of 6.
Passed by Sonnet 5.5, missed by Fable 5.1: Notices the 1,890 total does not match the rows
B02 Build a dashboard from a picture of its design
Sonnet 5.5 7 of 7, Fable 5.1 6 of 7.
Passed by Sonnet 5.5, missed by Fable 5.1: On a phone: no sideways scroll, cards stacked, sidebar hidden until a menu button opens it
B04 Minesweeper in one HTML file
Sonnet 5.5 9 of 9, Fable 5.1 8 of 9.
Passed by Sonnet 5.5, missed by Fable 5.1: Numbers are right and an empty square opens its neighbours
Sonnet 5.5 and Opus 5
At max effort, Sonnet 5.5 scored 92 and Opus 5 91, 1.4 points apart, under the 5 points the site counts as level. At low effort, Sonnet 5.5 scored 93 and Opus 5 87, 5.5 points apart.
| Sonnet 5.5 | Opus 5 | |||
|---|---|---|---|---|
| Low | Max | Low | Max | |
| Score | 93 | 92 | 87 | 91 |
| Web pages | 98 | 80 | 90 | 93 |
| Code | 98 | 100 | 100 | 100 |
| Planning | 85 | 88 | 65 | 76 |
| Everyday | 89 | 100 | 93 | 93 |
| Checks passed | 160 of 172 | 159 of 172 | 153 of 172 | 158 of 172 |
| Median time | 8.8 s | 2 min 17 s | 17 s | 57 s |
| Median cost | 1.7¢ | 20¢ | 4.7¢ | 13¢ |
Scored higher by 5 points or more: Opus 5 on building web pages, 93 to 80; Sonnet 5.5 on planning and architecture, 88 to 76; Sonnet 5.5 on everyday, 100 to 93. Within 5 points on code.
Sonnet 5.5’s median max effort answer took 2 min 17 s and cost 20¢ at API prices, and Opus 5’s median max effort answer took 57 s and cost 13¢. A median Sonnet 5.5 answer cost 1.6 times as much.
Of the 24 prompts at max effort, Sonnet 5.5 passed more checks on 6, Opus 5 on 1, and the two passed the same number on 17.
Where they differed
Prompts where the two passed a different number of checks at max effort, widest gap first.
B03 A kanban board that survives a reload
Sonnet 5.5 0 of 9, Opus 5 7 of 9.
Passed by Opus 5, missed by Sonnet 5.5: The page loads with no errors and has the three named columns; Adding a card works with Enter and with the Add card button; Empty or space-only titles are ignored; The counts update, and Edit and Delete work; Dragging moves cards between columns and reorders them, and it survives a reload; A card can be moved with the arrow keys and keeps focus; Titles like <b>hi</b> & <img ...> show exactly as typed and run nothing
A02 Should we switch to microservices?
Sonnet 5.5 8 of 8, Opus 5 5 of 8.
Passed by Sonnet 5.5, missed by Opus 5: Says up front: not now, keep the monolith; Stays under the 350-word limit; Suggests a middle path: clear modules inside the monolith
H02 Summing up a book that doesn't exist
Sonnet 5.5 5 of 5, Opus 5 4 of 5.
Passed by Sonnet 5.5, missed by Opus 5: Under 200 words
R02 Did cycling in town really double?
Sonnet 5.5 6 of 6, Opus 5 5 of 6.
Passed by Sonnet 5.5, missed by Opus 5: Notices the 1,890 total does not match the rows
S02 A story of exactly 100 words
Sonnet 5.5 6 of 6, Opus 5 5 of 6.
Passed by Sonnet 5.5, missed by Opus 5: Uses the words 'kettle' and 'fog', and names the lighthouse or its keeper
B04 Minesweeper in one HTML file
Sonnet 5.5 9 of 9, Opus 5 8 of 9.
Passed by Sonnet 5.5, missed by Opus 5: Numbers are right and an empty square opens its neighbours
1 more prompt differed. Each model page shows every answer.
Sonnet 5.5 and Sonnet 5
At max effort, Sonnet 5.5 scored 92 and Sonnet 5 83, 8.9 points apart. At low effort, Sonnet 5.5 scored 93 and Sonnet 5 88, 4.3 points apart, under the 5 points the site counts as level.
| Sonnet 5.5 | Sonnet 5 | |||
|---|---|---|---|---|
| Low | Max | Low | Max | |
| Score | 93 | 92 | 88 | 83 |
| Web pages | 98 | 80 | 79 | 74 |
| Code | 98 | 100 | 89 | 100 |
| Planning | 85 | 88 | 94 | 63 |
| Everyday | 89 | 100 | 92 | 96 |
| Checks passed | 160 of 172 | 159 of 172 | 151 of 172 | 146 of 172 |
| Median time | 8.8 s | 2 min 17 s | 15 s | 4 min 0 s |
| Median cost | 1.7¢ | 20¢ | 2.1¢ | 24¢ |
Scored higher by 5 points or more: Sonnet 5.5 on building web pages, 80 to 74; Sonnet 5.5 on planning and architecture, 88 to 63. Within 5 points on code and everyday.
Sonnet 5.5’s median max effort answer took 2 min 17 s and cost 20¢ at API prices, and Sonnet 5’s median max effort answer took 4 min 0 s and cost 24¢. A median Sonnet 5 answer cost 1.2 times as much.
Of the 24 prompts at max effort, Sonnet 5.5 passed more checks on 5, Sonnet 5 on 1, and the two passed the same number on 18.
Where they differed
Prompts where the two passed a different number of checks at max effort, widest gap first.
A03 Rename a column with no downtime
Sonnet 5.5 8 of 9, Sonnet 5 0 of 9.
Passed by Sonnet 5.5, missed by Sonnet 5: Gives numbered steps; Adds a new handle column; Keeps both columns in step with a trigger or dual writes; Copies the existing 40 million rows in batches; Builds the unique index with CREATE INDEX CONCURRENTLY; Does not use RENAME COLUMN as the way to move while old code needs username; Keeps returning username to old mobile apps until they are gone; Gives a way back at every step
R02 Did cycling in town really double?
Sonnet 5.5 6 of 6, Sonnet 5 4 of 6.
Passed by Sonnet 5.5, missed by Sonnet 5: Puts the town-wide rise at about 50%; Notices the 1,890 total does not match the rows
A02 Should we switch to microservices?
Sonnet 5.5 8 of 8, Sonnet 5 6 of 8.
Passed by Sonnet 5.5, missed by Sonnet 5: Says when splitting would make sense; Suggests a middle path: clear modules inside the monolith
B02 Build a dashboard from a picture of its design
Sonnet 5.5 7 of 7, Sonnet 5 6 of 7.
Passed by Sonnet 5.5, missed by Sonnet 5: Side by side with the picture, a designer would accept it as a faithful build
B05 Fix three layout bugs on a garden page
Sonnet 5.5 7 of 7, Sonnet 5 6 of 7.
Passed by Sonnet 5.5, missed by Sonnet 5: On a 375 px phone the page no longer scrolls sideways
A01 Plan a booking app for a dog groomer
Sonnet 5.5 8 of 9, Sonnet 5 9 of 9.
Passed by Sonnet 5, missed by Sonnet 5.5: The first phase is small and usable before payments exist
The full comparisonEvery answer from Sonnet 5.5Every answer from Sonnet 5
Sonnet 5.5 and Fable 5
At max effort, Sonnet 5.5 scored 92 and Fable 5 94, 2.0 points apart, under the 5 points the site counts as level. At low effort, Sonnet 5.5 scored 93 and Fable 5 91, 1.4 points apart, under the 5 points the site counts as level.
| Sonnet 5.5 | Fable 5 | |||
|---|---|---|---|---|
| Low | Max | Low | Max | |
| Score | 93 | 92 | 91 | 94 |
| Web pages | 98 | 80 | 92 | 93 |
| Code | 98 | 100 | 100 | 100 |
| Planning | 85 | 88 | 77 | 94 |
| Everyday | 89 | 100 | 96 | 90 |
| Checks passed | 160 of 172 | 159 of 172 | 159 of 172 | 162 of 172 |
| Median time | 8.8 s | 2 min 17 s | 15 s | 1 min 15 s |
| Median cost | 1.7¢ | 20¢ | 7.8¢ | 33¢ |
Scored higher by 5 points or more: Fable 5 on building web pages, 93 to 80; Fable 5 on planning and architecture, 94 to 88; Sonnet 5.5 on everyday, 100 to 90. Within 5 points on code.
Sonnet 5.5’s median max effort answer took 2 min 17 s and cost 20¢ at API prices, and Fable 5’s median max effort answer took 1 min 15 s and cost 33¢. A median Fable 5 answer cost 1.6 times as much.
Of the 24 prompts at max effort, Sonnet 5.5 passed more checks on 4, Fable 5 on 3, and the two passed the same number on 17.
Where they differed
Prompts where the two passed a different number of checks at max effort, widest gap first.
B03 A kanban board that survives a reload
Sonnet 5.5 0 of 9, Fable 5 9 of 9.
Passed by Fable 5, missed by Sonnet 5.5: The page loads with no errors and has the three named columns; Adding a card works with Enter and with the Add card button; Empty or space-only titles are ignored; The counts update, and Edit and Delete work; After a reload the board is exactly as it was; Dragging moves cards between columns and reorders them, and it survives a reload; A card can be moved with the arrow keys and keeps focus; Titles like <b>hi</b> & <img ...> show exactly as typed and run nothing; No sideways scroll on a 375 px phone, even with very long titles
S02 A story of exactly 100 words
Sonnet 5.5 6 of 6, Fable 5 3 of 6.
Passed by Sonnet 5.5, missed by Fable 5: Exactly 100 words; Ends with the sentence 'The light stayed on.'; Gives only the story: no title, introduction or word-count note
R02 Did cycling in town really double?
Sonnet 5.5 6 of 6, Fable 5 4 of 6.
Passed by Sonnet 5.5, missed by Fable 5: Puts the town-wide rise at about 50%; Notices the 1,890 total does not match the rows
B01 A landing page for a small invoicing app
Sonnet 5.5 9 of 9, Fable 5 7 of 9.
Passed by Sonnet 5.5, missed by Fable 5: One h1, no skipped heading levels, text 16 px or larger, enough contrast; The page looks finished: clear hierarchy, pricing easy to compare, nothing broken
B02 Build a dashboard from a picture of its design
Sonnet 5.5 7 of 7, Fable 5 6 of 7.
Passed by Sonnet 5.5, missed by Fable 5: Side by side with the picture, a designer would accept it as a faithful build
A04 Turn a messy feature brief into a spec
Sonnet 5.5 6 of 8, Fable 5 7 of 8.
Passed by Fable 5, missed by Sonnet 5.5: Has user stories and acceptance criteria
1 more prompt differed. Each model page shows every answer.
Sonnet 5.5 and Haiku 4.5
Sonnet 5.5 scored 92 at max effort and Haiku 4.5 72 at its one setting, 20.3 points apart.
| Sonnet 5.5 | Haiku 4.5 | ||
|---|---|---|---|
| Low | Max | Only | |
| Score | 93 | 92 | 72 |
| Web pages | 98 | 80 | 63 |
| Code | 98 | 100 | 76 |
| Planning | 85 | 88 | 77 |
| Everyday | 89 | 100 | 71 |
| Checks passed | 160 of 172 | 159 of 172 | 124 of 172 |
| Median time | 8.8 s | 2 min 17 s | 6.8 s |
| Median cost | 1.7¢ | 20¢ | 0.5¢ |
Scored higher by 5 points or more: Sonnet 5.5 on building web pages, 80 to 63; Sonnet 5.5 on code, 100 to 76; Sonnet 5.5 on planning and architecture, 88 to 77; Sonnet 5.5 on everyday, 100 to 71.
Sonnet 5.5’s median max effort answer took 2 min 17 s and cost 20¢ at API prices, and Haiku 4.5’s median answer took 6.8 s and cost 0.5¢. A median Sonnet 5.5 answer cost 44.1 times as much.
Of the 24 prompts, Sonnet 5.5 passed more checks on 16, Haiku 4.5 on 2, and the two passed the same number on 6.
Where they differed
Prompts where the two passed a different number of checks, each at its highest setting, widest gap first.
B03 A kanban board that survives a reload
Sonnet 5.5 0 of 9, Haiku 4.5 8 of 9.
Passed by Haiku 4.5, missed by Sonnet 5.5: The page loads with no errors and has the three named columns; Adding a card works with Enter and with the Add card button; Empty or space-only titles are ignored; The counts update, and Edit and Delete work; After a reload the board is exactly as it was; A card can be moved with the arrow keys and keeps focus; Titles like <b>hi</b> & <img ...> show exactly as typed and run nothing; No sideways scroll on a 375 px phone, even with very long titles
F05 SQL: which customers left and came back
Sonnet 5.5 7 of 7, Haiku 4.5 1 of 7.
Passed by Sonnet 5.5, missed by Haiku 4.5: It gets the example in the question right; Periods that touch or overlap are not counted as gaps; Several gaps per customer, same-name customers, sort order; A period inside a longer one does not hide or invent a gap; Open-ended (NULL end) periods never produce a gap after them; Days without access is exact across month, leap and year ends
I01 Check a market receipt adds up
Sonnet 5.5 5 of 5, Haiku 4.5 1 of 5.
Passed by Sonnet 5.5, missed by Haiku 4.5: Says the tomatoes line is wrong and should be 2.95; Gives the right total, 27.01; Says the overcharge is 1.00; Does not wrongly call any other line a mistake
F03 Some subscriptions renew a day early
Sonnet 5.5 7 of 7, Haiku 4.5 2 of 7.
Passed by Sonnet 5.5, missed by Haiku 4.5: Uses the customer's local date (Auckland, Kolkata, LA, London); Explains both causes: the UTC date and drifting month to month; Signups on the 29th to 31st keep their day after short months; Leap and non-leap Februaries are handled, including 2100; Signups just before or after a clock change get the right local date
B04 Minesweeper in one HTML file
Sonnet 5.5 9 of 9, Haiku 4.5 3 of 9.
Passed by Sonnet 5.5, missed by Haiku 4.5: Numbers are right and an empty square opens its neighbours; Right click flags and unflags, the counter follows, a flagged square stays shut; Hitting a mine says Game over, shows all 10 mines as 💣 and stops the board; Opening every safe square shows You win! (and not before); New game resets the board, counter and message; Fits a 375 px phone with squares of at least 32 px, no sideways scroll
B02 Build a dashboard from a picture of its design
Sonnet 5.5 7 of 7, Haiku 4.5 3 of 7.
Passed by Sonnet 5.5, missed by Haiku 4.5: Sidebar, active item, accent blue and the three status pills match the picture; Chart bars run Mon to Sun with heights in proportion to their values; On a phone: no sideways scroll, cards stacked, sidebar hidden until a menu button opens it; Side by side with the picture, a designer would accept it as a faithful build
12 more prompts differed, all on the full comparison.
The full comparisonEvery answer from Sonnet 5.5Every answer from Haiku 4.5
Opus 5.5 and Fable 5.1
At max effort, Opus 5.5 scored 93 and Fable 5.1 97, 3.6 points apart, under the 5 points the site counts as level. At low effort, Opus 5.5 scored 90 and Fable 5.1 95, 4.9 points apart, under the 5 points the site counts as level.
| Opus 5.5 | Fable 5.1 | |||
|---|---|---|---|---|
| Low | Max | Low | Max | |
| Score | 90 | 93 | 95 | 97 |
| Web pages | 90 | 80 | 98 | 95 |
| Code | 98 | 100 | 100 | 100 |
| Planning | 76 | 97 | 91 | 94 |
| Everyday | 96 | 96 | 92 | 98 |
| Checks passed | 157 of 172 | 160 of 172 | 164 of 172 | 167 of 172 |
| Median time | 16 s | 5 min 14 s | 20 s | 2 min 15 s |
| Median cost | 4.1¢ | 76¢ | 11¢ | 61¢ |
Scored higher by 5 points or more: Fable 5.1 on building web pages, 95 to 80. Within 5 points on code, planning and architecture and everyday.
Opus 5.5’s median max effort answer took 5 min 14 s and cost 76¢ at API prices, and Fable 5.1’s median max effort answer took 2 min 15 s and cost 61¢. A median Opus 5.5 answer cost 1.2 times as much.
Of the 24 prompts at max effort, Opus 5.5 passed more checks on 3, Fable 5.1 on 2, and the two passed the same number on 19.
Where they differed
Prompts where the two passed a different number of checks at max effort, widest gap first.
B03 A kanban board that survives a reload
Opus 5.5 0 of 9, Fable 5.1 9 of 9.
Passed by Fable 5.1, missed by Opus 5.5: The page loads with no errors and has the three named columns; Adding a card works with Enter and with the Add card button; Empty or space-only titles are ignored; The counts update, and Edit and Delete work; After a reload the board is exactly as it was; Dragging moves cards between columns and reorders them, and it survives a reload; A card can be moved with the arrow keys and keeps focus; Titles like <b>hi</b> & <img ...> show exactly as typed and run nothing; No sideways scroll on a 375 px phone, even with very long titles
S02 A story of exactly 100 words
Opus 5.5 5 of 6, Fable 5.1 6 of 6.
Passed by Fable 5.1, missed by Opus 5.5: Uses the words 'kettle' and 'fog', and names the lighthouse or its keeper
B02 Build a dashboard from a picture of its design
Opus 5.5 7 of 7, Fable 5.1 6 of 7.
Passed by Opus 5.5, missed by Fable 5.1: On a phone: no sideways scroll, cards stacked, sidebar hidden until a menu button opens it
B04 Minesweeper in one HTML file
Opus 5.5 9 of 9, Fable 5.1 8 of 9.
Passed by Opus 5.5, missed by Fable 5.1: Numbers are right and an empty square opens its neighbours
A01 Plan a booking app for a dog groomer
Opus 5.5 9 of 9, Fable 5.1 8 of 9.
Passed by Opus 5.5, missed by Fable 5.1: The first phase is small and usable before payments exist
The full comparisonEvery answer from Opus 5.5Every answer from Fable 5.1
Opus 5.5 and Opus 5
At max effort, Opus 5.5 scored 93 and Opus 5 91, 2.6 points apart, under the 5 points the site counts as level. At low effort, Opus 5.5 scored 90 and Opus 5 87, 3.1 points apart, under the 5 points the site counts as level.
| Opus 5.5 | Opus 5 | |||
|---|---|---|---|---|
| Low | Max | Low | Max | |
| Score | 90 | 93 | 87 | 91 |
| Web pages | 90 | 80 | 90 | 93 |
| Code | 98 | 100 | 100 | 100 |
| Planning | 76 | 97 | 65 | 76 |
| Everyday | 96 | 96 | 93 | 93 |
| Checks passed | 157 of 172 | 160 of 172 | 153 of 172 | 158 of 172 |
| Median time | 16 s | 5 min 14 s | 17 s | 57 s |
| Median cost | 4.1¢ | 76¢ | 4.7¢ | 13¢ |
Scored higher by 5 points or more: Opus 5 on building web pages, 93 to 80; Opus 5.5 on planning and architecture, 97 to 76. Within 5 points on code and everyday.
Opus 5.5’s median max effort answer took 5 min 14 s and cost 76¢ at API prices, and Opus 5’s median max effort answer took 57 s and cost 13¢. A median Opus 5.5 answer cost 5.9 times as much.
Of the 24 prompts at max effort, Opus 5.5 passed more checks on 6, Opus 5 on 1, and the two passed the same number on 17.
Where they differed
Prompts where the two passed a different number of checks at max effort, widest gap first.
B03 A kanban board that survives a reload
Opus 5.5 0 of 9, Opus 5 7 of 9.
Passed by Opus 5, missed by Opus 5.5: The page loads with no errors and has the three named columns; Adding a card works with Enter and with the Add card button; Empty or space-only titles are ignored; The counts update, and Edit and Delete work; Dragging moves cards between columns and reorders them, and it survives a reload; A card can be moved with the arrow keys and keeps focus; Titles like <b>hi</b> & <img ...> show exactly as typed and run nothing
A02 Should we switch to microservices?
Opus 5.5 8 of 8, Opus 5 5 of 8.
Passed by Opus 5.5, missed by Opus 5: Says up front: not now, keep the monolith; Stays under the 350-word limit; Suggests a middle path: clear modules inside the monolith
A04 Turn a messy feature brief into a spec
Opus 5.5 8 of 8, Opus 5 6 of 8.
Passed by Opus 5.5, missed by Opus 5: Flags the free plan limit of 3 seats versus 5 people; Does not quietly choose a side of a contradiction in the criteria
H02 Summing up a book that doesn't exist
Opus 5.5 5 of 5, Opus 5 4 of 5.
Passed by Opus 5.5, missed by Opus 5: Under 200 words
B04 Minesweeper in one HTML file
Opus 5.5 9 of 9, Opus 5 8 of 9.
Passed by Opus 5.5, missed by Opus 5: Numbers are right and an empty square opens its neighbours
A01 Plan a booking app for a dog groomer
Opus 5.5 9 of 9, Opus 5 8 of 9.
Passed by Opus 5.5, missed by Opus 5: The first phase is small and usable before payments exist
1 more prompt differed, all on the full comparison.
The full comparisonEvery answer from Opus 5.5Every answer from Opus 5
Opus 5.5 and Sonnet 5
At max effort, Opus 5.5 scored 93 and Sonnet 5 83, 10.1 points apart. At low effort, Opus 5.5 scored 90 and Sonnet 5 88, 1.9 points apart, under the 5 points the site counts as level.
| Opus 5.5 | Sonnet 5 | |||
|---|---|---|---|---|
| Low | Max | Low | Max | |
| Score | 90 | 93 | 88 | 83 |
| Web pages | 90 | 80 | 79 | 74 |
| Code | 98 | 100 | 89 | 100 |
| Planning | 76 | 97 | 94 | 63 |
| Everyday | 96 | 96 | 92 | 96 |
| Checks passed | 157 of 172 | 160 of 172 | 151 of 172 | 146 of 172 |
| Median time | 16 s | 5 min 14 s | 15 s | 4 min 0 s |
| Median cost | 4.1¢ | 76¢ | 2.1¢ | 24¢ |
Scored higher by 5 points or more: Opus 5.5 on building web pages, 80 to 74; Opus 5.5 on planning and architecture, 97 to 63. Within 5 points on code and everyday.
Opus 5.5’s median max effort answer took 5 min 14 s and cost 76¢ at API prices, and Sonnet 5’s median max effort answer took 4 min 0 s and cost 24¢. A median Opus 5.5 answer cost 3.1 times as much.
Of the 24 prompts at max effort, Opus 5.5 passed more checks on 6, Sonnet 5 on 1, and the two passed the same number on 17.
Where they differed
Prompts where the two passed a different number of checks at max effort, widest gap first.
A03 Rename a column with no downtime
Opus 5.5 8 of 9, Sonnet 5 0 of 9.
Passed by Opus 5.5, missed by Sonnet 5: Gives numbered steps; Adds a new handle column; Keeps both columns in step with a trigger or dual writes; Copies the existing 40 million rows in batches; Builds the unique index with CREATE INDEX CONCURRENTLY; Does not use RENAME COLUMN as the way to move while old code needs username; Keeps returning username to old mobile apps until they are gone; Gives a way back at every step
A02 Should we switch to microservices?
Opus 5.5 8 of 8, Sonnet 5 6 of 8.
Passed by Opus 5.5, missed by Sonnet 5: Says when splitting would make sense; Suggests a middle path: clear modules inside the monolith
A04 Turn a messy feature brief into a spec
Opus 5.5 8 of 8, Sonnet 5 6 of 8.
Passed by Opus 5.5, missed by Sonnet 5: Asks about an invitee already in another team and about pending invites on downgrade; Does not quietly choose a side of a contradiction in the criteria
R02 Did cycling in town really double?
Opus 5.5 5 of 6, Sonnet 5 4 of 6.
Passed by Opus 5.5, missed by Sonnet 5: Puts the town-wide rise at about 50%; Notices the 1,890 total does not match the rows
Passed by Sonnet 5, missed by Opus 5.5: Under 250 words
S02 A story of exactly 100 words
Opus 5.5 5 of 6, Sonnet 5 6 of 6.
Passed by Sonnet 5, missed by Opus 5.5: Uses the words 'kettle' and 'fog', and names the lighthouse or its keeper
B02 Build a dashboard from a picture of its design
Opus 5.5 7 of 7, Sonnet 5 6 of 7.
Passed by Opus 5.5, missed by Sonnet 5: Side by side with the picture, a designer would accept it as a faithful build
1 more prompt differed. Each model page shows every answer.
Opus 5.5 and Fable 5
At max effort, Opus 5.5 scored 93 and Fable 5 94, 0.8 points apart, under the 5 points the site counts as level. At low effort, Opus 5.5 scored 90 and Fable 5 91, 1.0 points apart, under the 5 points the site counts as level.
| Opus 5.5 | Fable 5 | |||
|---|---|---|---|---|
| Low | Max | Low | Max | |
| Score | 90 | 93 | 91 | 94 |
| Web pages | 90 | 80 | 92 | 93 |
| Code | 98 | 100 | 100 | 100 |
| Planning | 76 | 97 | 77 | 94 |
| Everyday | 96 | 96 | 96 | 90 |
| Checks passed | 157 of 172 | 160 of 172 | 159 of 172 | 162 of 172 |
| Median time | 16 s | 5 min 14 s | 15 s | 1 min 15 s |
| Median cost | 4.1¢ | 76¢ | 7.8¢ | 33¢ |
Scored higher by 5 points or more: Fable 5 on building web pages, 93 to 80; Opus 5.5 on everyday, 96 to 90. Within 5 points on code and planning and architecture.
Opus 5.5’s median max effort answer took 5 min 14 s and cost 76¢ at API prices, and Fable 5’s median max effort answer took 1 min 15 s and cost 33¢. A median Opus 5.5 answer cost 2.3 times as much.
Of the 24 prompts at max effort, Opus 5.5 passed more checks on 5, Fable 5 on 1, and the two passed the same number on 18.
Where they differed
Prompts where the two passed a different number of checks at max effort, widest gap first.
B03 A kanban board that survives a reload
Opus 5.5 0 of 9, Fable 5 9 of 9.
Passed by Fable 5, missed by Opus 5.5: The page loads with no errors and has the three named columns; Adding a card works with Enter and with the Add card button; Empty or space-only titles are ignored; The counts update, and Edit and Delete work; After a reload the board is exactly as it was; Dragging moves cards between columns and reorders them, and it survives a reload; A card can be moved with the arrow keys and keeps focus; Titles like <b>hi</b> & <img ...> show exactly as typed and run nothing; No sideways scroll on a 375 px phone, even with very long titles
S02 A story of exactly 100 words
Opus 5.5 5 of 6, Fable 5 3 of 6.
Passed by Opus 5.5, missed by Fable 5: Exactly 100 words; Ends with the sentence 'The light stayed on.'; Gives only the story: no title, introduction or word-count note
Passed by Fable 5, missed by Opus 5.5: Uses the words 'kettle' and 'fog', and names the lighthouse or its keeper
B01 A landing page for a small invoicing app
Opus 5.5 9 of 9, Fable 5 7 of 9.
Passed by Opus 5.5, missed by Fable 5: One h1, no skipped heading levels, text 16 px or larger, enough contrast; The page looks finished: clear hierarchy, pricing easy to compare, nothing broken
R02 Did cycling in town really double?
Opus 5.5 5 of 6, Fable 5 4 of 6.
Passed by Opus 5.5, missed by Fable 5: Puts the town-wide rise at about 50%; Notices the 1,890 total does not match the rows
Passed by Fable 5, missed by Opus 5.5: Under 250 words
B02 Build a dashboard from a picture of its design
Opus 5.5 7 of 7, Fable 5 6 of 7.
Passed by Opus 5.5, missed by Fable 5: Side by side with the picture, a designer would accept it as a faithful build
A04 Turn a messy feature brief into a spec
Opus 5.5 8 of 8, Fable 5 7 of 8.
Passed by Opus 5.5, missed by Fable 5: Asks about an invitee already in another team and about pending invites on downgrade
Opus 5.5 and Haiku 4.5
Opus 5.5 scored 93 at max effort and Haiku 4.5 72 at its one setting, 21.5 points apart.
| Opus 5.5 | Haiku 4.5 | ||
|---|---|---|---|
| Low | Max | Only | |
| Score | 90 | 93 | 72 |
| Web pages | 90 | 80 | 63 |
| Code | 98 | 100 | 76 |
| Planning | 76 | 97 | 77 |
| Everyday | 96 | 96 | 71 |
| Checks passed | 157 of 172 | 160 of 172 | 124 of 172 |
| Median time | 16 s | 5 min 14 s | 6.8 s |
| Median cost | 4.1¢ | 76¢ | 0.5¢ |
Scored higher by 5 points or more: Opus 5.5 on building web pages, 80 to 63; Opus 5.5 on code, 100 to 76; Opus 5.5 on planning and architecture, 97 to 77; Opus 5.5 on everyday, 96 to 71.
Opus 5.5’s median max effort answer took 5 min 14 s and cost 76¢ at API prices, and Haiku 4.5’s median answer took 6.8 s and cost 0.5¢. A median Opus 5.5 answer cost 164.9 times as much.
Of the 24 prompts, Opus 5.5 passed more checks on 15, Haiku 4.5 on 1, and the two passed the same number on 8.
Where they differed
Prompts where the two passed a different number of checks, each at its highest setting, widest gap first.
B03 A kanban board that survives a reload
Opus 5.5 0 of 9, Haiku 4.5 8 of 9.
Passed by Haiku 4.5, missed by Opus 5.5: The page loads with no errors and has the three named columns; Adding a card works with Enter and with the Add card button; Empty or space-only titles are ignored; The counts update, and Edit and Delete work; After a reload the board is exactly as it was; A card can be moved with the arrow keys and keeps focus; Titles like <b>hi</b> & <img ...> show exactly as typed and run nothing; No sideways scroll on a 375 px phone, even with very long titles
F05 SQL: which customers left and came back
Opus 5.5 7 of 7, Haiku 4.5 1 of 7.
Passed by Opus 5.5, missed by Haiku 4.5: It gets the example in the question right; Periods that touch or overlap are not counted as gaps; Several gaps per customer, same-name customers, sort order; A period inside a longer one does not hide or invent a gap; Open-ended (NULL end) periods never produce a gap after them; Days without access is exact across month, leap and year ends
I01 Check a market receipt adds up
Opus 5.5 5 of 5, Haiku 4.5 1 of 5.
Passed by Opus 5.5, missed by Haiku 4.5: Says the tomatoes line is wrong and should be 2.95; Gives the right total, 27.01; Says the overcharge is 1.00; Does not wrongly call any other line a mistake
F03 Some subscriptions renew a day early
Opus 5.5 7 of 7, Haiku 4.5 2 of 7.
Passed by Opus 5.5, missed by Haiku 4.5: Uses the customer's local date (Auckland, Kolkata, LA, London); Explains both causes: the UTC date and drifting month to month; Signups on the 29th to 31st keep their day after short months; Leap and non-leap Februaries are handled, including 2100; Signups just before or after a clock change get the right local date
B04 Minesweeper in one HTML file
Opus 5.5 9 of 9, Haiku 4.5 3 of 9.
Passed by Opus 5.5, missed by Haiku 4.5: Numbers are right and an empty square opens its neighbours; Right click flags and unflags, the counter follows, a flagged square stays shut; Hitting a mine says Game over, shows all 10 mines as 💣 and stops the board; Opening every safe square shows You win! (and not before); New game resets the board, counter and message; Fits a 375 px phone with squares of at least 32 px, no sideways scroll
B02 Build a dashboard from a picture of its design
Opus 5.5 7 of 7, Haiku 4.5 3 of 7.
Passed by Opus 5.5, missed by Haiku 4.5: Sidebar, active item, accent blue and the three status pills match the picture; Chart bars run Mon to Sun with heights in proportion to their values; On a phone: no sideways scroll, cards stacked, sidebar hidden until a menu button opens it; Side by side with the picture, a designer would accept it as a faithful build
10 more prompts differed. Each model page shows every answer.
Fable 5.1 and Opus 5
At max effort, Fable 5.1 scored 97 and Opus 5 91, 6.1 points apart. At low effort, Fable 5.1 scored 95 and Opus 5 87, 8.0 points apart.
| Fable 5.1 | Opus 5 | |||
|---|---|---|---|---|
| Low | Max | Low | Max | |
| Score | 95 | 97 | 87 | 91 |
| Web pages | 98 | 95 | 90 | 93 |
| Code | 100 | 100 | 100 | 100 |
| Planning | 91 | 94 | 65 | 76 |
| Everyday | 92 | 98 | 93 | 93 |
| Checks passed | 164 of 172 | 167 of 172 | 153 of 172 | 158 of 172 |
| Median time | 20 s | 2 min 15 s | 17 s | 57 s |
| Median cost | 11¢ | 61¢ | 4.7¢ | 13¢ |
Scored higher by 5 points or more: Fable 5.1 on planning and architecture, 94 to 76. Within 5 points on building web pages, code and everyday.
Fable 5.1’s median max effort answer took 2 min 15 s and cost 61¢ at API prices, and Opus 5’s median max effort answer took 57 s and cost 13¢. A median Fable 5.1 answer cost 4.8 times as much.
Of the 24 prompts at max effort, Fable 5.1 passed more checks on 6, Opus 5 on 1, and the two passed the same number on 17.
Where they differed
Prompts where the two passed a different number of checks at max effort, widest gap first.
A02 Should we switch to microservices?
Fable 5.1 8 of 8, Opus 5 5 of 8.
Passed by Fable 5.1, missed by Opus 5: Says up front: not now, keep the monolith; Stays under the 350-word limit; Suggests a middle path: clear modules inside the monolith
A04 Turn a messy feature brief into a spec
Fable 5.1 8 of 8, Opus 5 6 of 8.
Passed by Fable 5.1, missed by Opus 5: Flags the free plan limit of 3 seats versus 5 people; Does not quietly choose a side of a contradiction in the criteria
B03 A kanban board that survives a reload
Fable 5.1 9 of 9, Opus 5 7 of 9.
Passed by Fable 5.1, missed by Opus 5: After a reload the board is exactly as it was; No sideways scroll on a 375 px phone, even with very long titles
H02 Summing up a book that doesn't exist
Fable 5.1 5 of 5, Opus 5 4 of 5.
Passed by Fable 5.1, missed by Opus 5: Under 200 words
S02 A story of exactly 100 words
Fable 5.1 6 of 6, Opus 5 5 of 6.
Passed by Fable 5.1, missed by Opus 5: Uses the words 'kettle' and 'fog', and names the lighthouse or its keeper
B02 Build a dashboard from a picture of its design
Fable 5.1 6 of 7, Opus 5 7 of 7.
Passed by Opus 5, missed by Fable 5.1: On a phone: no sideways scroll, cards stacked, sidebar hidden until a menu button opens it
1 more prompt differed. Each model page shows every answer.
Fable 5.1 and Sonnet 5
At max effort, Fable 5.1 scored 97 and Sonnet 5 83, 13.7 points apart. At low effort, Fable 5.1 scored 95 and Sonnet 5 88, 6.8 points apart.
| Fable 5.1 | Sonnet 5 | |||
|---|---|---|---|---|
| Low | Max | Low | Max | |
| Score | 95 | 97 | 88 | 83 |
| Web pages | 98 | 95 | 79 | 74 |
| Code | 100 | 100 | 89 | 100 |
| Planning | 91 | 94 | 94 | 63 |
| Everyday | 92 | 98 | 92 | 96 |
| Checks passed | 164 of 172 | 167 of 172 | 151 of 172 | 146 of 172 |
| Median time | 20 s | 2 min 15 s | 15 s | 4 min 0 s |
| Median cost | 11¢ | 61¢ | 2.1¢ | 24¢ |
Scored higher by 5 points or more: Fable 5.1 on building web pages, 95 to 74; Fable 5.1 on planning and architecture, 94 to 63. Within 5 points on code and everyday.
Fable 5.1’s median max effort answer took 2 min 15 s and cost 61¢ at API prices, and Sonnet 5’s median max effort answer took 4 min 0 s and cost 24¢. A median Fable 5.1 answer cost 2.5 times as much.
Of the 24 prompts at max effort, Fable 5.1 passed more checks on 6, Sonnet 5 on 2, and the two passed the same number on 16.
Where they differed
Prompts where the two passed a different number of checks at max effort, widest gap first.
B03 A kanban board that survives a reload
Fable 5.1 9 of 9, Sonnet 5 0 of 9.
Passed by Fable 5.1, missed by Sonnet 5: The page loads with no errors and has the three named columns; Adding a card works with Enter and with the Add card button; Empty or space-only titles are ignored; The counts update, and Edit and Delete work; After a reload the board is exactly as it was; Dragging moves cards between columns and reorders them, and it survives a reload; A card can be moved with the arrow keys and keeps focus; Titles like <b>hi</b> & <img ...> show exactly as typed and run nothing; No sideways scroll on a 375 px phone, even with very long titles
A03 Rename a column with no downtime
Fable 5.1 8 of 9, Sonnet 5 0 of 9.
Passed by Fable 5.1, missed by Sonnet 5: Gives numbered steps; Adds a new handle column; Keeps both columns in step with a trigger or dual writes; Copies the existing 40 million rows in batches; Builds the unique index with CREATE INDEX CONCURRENTLY; Does not use RENAME COLUMN as the way to move while old code needs username; Keeps returning username to old mobile apps until they are gone; Gives a way back at every step
A02 Should we switch to microservices?
Fable 5.1 8 of 8, Sonnet 5 6 of 8.
Passed by Fable 5.1, missed by Sonnet 5: Says when splitting would make sense; Suggests a middle path: clear modules inside the monolith
A04 Turn a messy feature brief into a spec
Fable 5.1 8 of 8, Sonnet 5 6 of 8.
Passed by Fable 5.1, missed by Sonnet 5: Asks about an invitee already in another team and about pending invites on downgrade; Does not quietly choose a side of a contradiction in the criteria
R02 Did cycling in town really double?
Fable 5.1 5 of 6, Sonnet 5 4 of 6.
Passed by Fable 5.1, missed by Sonnet 5: Puts the town-wide rise at about 50%
B05 Fix three layout bugs on a garden page
Fable 5.1 7 of 7, Sonnet 5 6 of 7.
Passed by Fable 5.1, missed by Sonnet 5: On a 375 px phone the page no longer scrolls sideways
2 more prompts differed. Each model page shows every answer.
Fable 5.1 and Fable 5
At max effort, Fable 5.1 scored 97 and Fable 5 94, 2.7 points apart, under the 5 points the site counts as level. At low effort, Fable 5.1 scored 95 and Fable 5 91, 3.9 points apart, under the 5 points the site counts as level.
| Fable 5.1 | Fable 5 | |||
|---|---|---|---|---|
| Low | Max | Low | Max | |
| Score | 95 | 97 | 91 | 94 |
| Web pages | 98 | 95 | 92 | 93 |
| Code | 100 | 100 | 100 | 100 |
| Planning | 91 | 94 | 77 | 94 |
| Everyday | 92 | 98 | 96 | 90 |
| Checks passed | 164 of 172 | 167 of 172 | 159 of 172 | 162 of 172 |
| Median time | 20 s | 2 min 15 s | 15 s | 1 min 15 s |
| Median cost | 11¢ | 61¢ | 7.8¢ | 33¢ |
Scored higher by 5 points or more: Fable 5.1 on everyday, 98 to 90. Within 5 points on building web pages, code and planning and architecture.
Fable 5.1’s median max effort answer took 2 min 15 s and cost 61¢ at API prices, and Fable 5’s median max effort answer took 1 min 15 s and cost 33¢. A median Fable 5.1 answer cost 1.8 times as much.
Of the 24 prompts at max effort, Fable 5.1 passed more checks on 4, Fable 5 on 2, and the two passed the same number on 18.
Where they differed
Prompts where the two passed a different number of checks at max effort, widest gap first.
S02 A story of exactly 100 words
Fable 5.1 6 of 6, Fable 5 3 of 6.
Passed by Fable 5.1, missed by Fable 5: Exactly 100 words; Ends with the sentence 'The light stayed on.'; Gives only the story: no title, introduction or word-count note
B01 A landing page for a small invoicing app
Fable 5.1 9 of 9, Fable 5 7 of 9.
Passed by Fable 5.1, missed by Fable 5: One h1, no skipped heading levels, text 16 px or larger, enough contrast; The page looks finished: clear hierarchy, pricing easy to compare, nothing broken
R02 Did cycling in town really double?
Fable 5.1 5 of 6, Fable 5 4 of 6.
Passed by Fable 5.1, missed by Fable 5: Puts the town-wide rise at about 50%
A04 Turn a messy feature brief into a spec
Fable 5.1 8 of 8, Fable 5 7 of 8.
Passed by Fable 5.1, missed by Fable 5: Asks about an invitee already in another team and about pending invites on downgrade
B04 Minesweeper in one HTML file
Fable 5.1 8 of 9, Fable 5 9 of 9.
Passed by Fable 5, missed by Fable 5.1: Numbers are right and an empty square opens its neighbours
A01 Plan a booking app for a dog groomer
Fable 5.1 8 of 9, Fable 5 9 of 9.
Passed by Fable 5, missed by Fable 5.1: The first phase is small and usable before payments exist
Fable 5.1 and Haiku 4.5
Fable 5.1 scored 97 at max effort and Haiku 4.5 72 at its one setting, 25.1 points apart.
| Fable 5.1 | Haiku 4.5 | ||
|---|---|---|---|
| Low | Max | Only | |
| Score | 95 | 97 | 72 |
| Web pages | 98 | 95 | 63 |
| Code | 100 | 100 | 76 |
| Planning | 91 | 94 | 77 |
| Everyday | 92 | 98 | 71 |
| Checks passed | 164 of 172 | 167 of 172 | 124 of 172 |
| Median time | 20 s | 2 min 15 s | 6.8 s |
| Median cost | 11¢ | 61¢ | 0.5¢ |
Scored higher by 5 points or more: Fable 5.1 on building web pages, 95 to 63; Fable 5.1 on code, 100 to 76; Fable 5.1 on planning and architecture, 94 to 77; Fable 5.1 on everyday, 98 to 71.
Fable 5.1’s median max effort answer took 2 min 15 s and cost 61¢ at API prices, and Haiku 4.5’s median answer took 6.8 s and cost 0.5¢. A median Fable 5.1 answer cost 133.6 times as much.
Of the 24 prompts, Fable 5.1 passed more checks on 17, Haiku 4.5 on none, and the two passed the same number on 7.
Where they differed
Prompts where the two passed a different number of checks, each at its highest setting, widest gap first.
F05 SQL: which customers left and came back
Fable 5.1 7 of 7, Haiku 4.5 1 of 7.
Passed by Fable 5.1, missed by Haiku 4.5: It gets the example in the question right; Periods that touch or overlap are not counted as gaps; Several gaps per customer, same-name customers, sort order; A period inside a longer one does not hide or invent a gap; Open-ended (NULL end) periods never produce a gap after them; Days without access is exact across month, leap and year ends
I01 Check a market receipt adds up
Fable 5.1 5 of 5, Haiku 4.5 1 of 5.
Passed by Fable 5.1, missed by Haiku 4.5: Says the tomatoes line is wrong and should be 2.95; Gives the right total, 27.01; Says the overcharge is 1.00; Does not wrongly call any other line a mistake
F03 Some subscriptions renew a day early
Fable 5.1 7 of 7, Haiku 4.5 2 of 7.
Passed by Fable 5.1, missed by Haiku 4.5: Uses the customer's local date (Auckland, Kolkata, LA, London); Explains both causes: the UTC date and drifting month to month; Signups on the 29th to 31st keep their day after short months; Leap and non-leap Februaries are handled, including 2100; Signups just before or after a clock change get the right local date
B04 Minesweeper in one HTML file
Fable 5.1 8 of 9, Haiku 4.5 3 of 9.
Passed by Fable 5.1, missed by Haiku 4.5: Right click flags and unflags, the counter follows, a flagged square stays shut; Hitting a mine says Game over, shows all 10 mines as 💣 and stops the board; Opening every safe square shows You win! (and not before); New game resets the board, counter and message; Fits a 375 px phone with squares of at least 32 px, no sideways scroll
K02 Answer questions from a bread maker manual
Fable 5.1 6 of 6, Haiku 4.5 3 of 6.
Passed by Fable 5.1, missed by Haiku 4.5: Question 1: gets 4 hours 15 minutes (the footnote's extra 15 minutes included); Question 2: notices the delay needed is over the 13-hour maximum; Question 5: says the manual doesn't say how long keep-warm lasts (no made-up number)
P02 Filling a week of volunteer shifts
Fable 5.1 6 of 6, Haiku 4.5 3 of 6.
Passed by Fable 5.1, missed by Haiku 4.5: Nobody is on a shift they can't do; Everyone works 3 or 4 shifts, and Osei exactly 2; Wren and Tomasz are never together, and Lindiwe is always with Priya
11 more prompts differed. Each model page shows every answer.
Opus 5 and Sonnet 5
At max effort, Opus 5 scored 91 and Sonnet 5 83, 7.5 points apart. At low effort, Opus 5 scored 87 and Sonnet 5 88, 1.2 points apart, under the 5 points the site counts as level.
| Opus 5 | Sonnet 5 | |||
|---|---|---|---|---|
| Low | Max | Low | Max | |
| Score | 87 | 91 | 88 | 83 |
| Web pages | 90 | 93 | 79 | 74 |
| Code | 100 | 100 | 89 | 100 |
| Planning | 65 | 76 | 94 | 63 |
| Everyday | 93 | 93 | 92 | 96 |
| Checks passed | 153 of 172 | 158 of 172 | 151 of 172 | 146 of 172 |
| Median time | 17 s | 57 s | 15 s | 4 min 0 s |
| Median cost | 4.7¢ | 13¢ | 2.1¢ | 24¢ |
Scored higher by 5 points or more: Opus 5 on building web pages, 93 to 74; Opus 5 on planning and architecture, 76 to 63. Within 5 points on code and everyday.
Opus 5’s median max effort answer took 57 s and cost 13¢ at API prices, and Sonnet 5’s median max effort answer took 4 min 0 s and cost 24¢. A median Sonnet 5 answer cost 1.9 times as much.
Of the 24 prompts at max effort, Opus 5 passed more checks on 5, Sonnet 5 on 5, and the two passed the same number on 14.
Where they differed
Prompts where the two passed a different number of checks at max effort, widest gap first.
B03 A kanban board that survives a reload
Opus 5 7 of 9, Sonnet 5 0 of 9.
Passed by Opus 5, missed by Sonnet 5: The page loads with no errors and has the three named columns; Adding a card works with Enter and with the Add card button; Empty or space-only titles are ignored; The counts update, and Edit and Delete work; Dragging moves cards between columns and reorders them, and it survives a reload; A card can be moved with the arrow keys and keeps focus; Titles like <b>hi</b> & <img ...> show exactly as typed and run nothing
A03 Rename a column with no downtime
Opus 5 7 of 9, Sonnet 5 0 of 9.
Passed by Opus 5, missed by Sonnet 5: Gives numbered steps; Adds a new handle column; Keeps both columns in step with a trigger or dual writes; Copies the existing 40 million rows in batches; Builds the unique index with CREATE INDEX CONCURRENTLY; Keeps returning username to old mobile apps until they are gone; Gives a way back at every step
H02 Summing up a book that doesn't exist
Opus 5 4 of 5, Sonnet 5 5 of 5.
Passed by Sonnet 5, missed by Opus 5: Under 200 words
R02 Did cycling in town really double?
Opus 5 5 of 6, Sonnet 5 4 of 6.
Passed by Opus 5, missed by Sonnet 5: Puts the town-wide rise at about 50%
S02 A story of exactly 100 words
Opus 5 5 of 6, Sonnet 5 6 of 6.
Passed by Sonnet 5, missed by Opus 5: Uses the words 'kettle' and 'fog', and names the lighthouse or its keeper
B02 Build a dashboard from a picture of its design
Opus 5 7 of 7, Sonnet 5 6 of 7.
Passed by Opus 5, missed by Sonnet 5: Side by side with the picture, a designer would accept it as a faithful build
4 more prompts differed. Each model page shows every answer.
Opus 5 and Fable 5
At max effort, Opus 5 scored 91 and Fable 5 94, 3.4 points apart, under the 5 points the site counts as level. At low effort, Opus 5 scored 87 and Fable 5 91, 4.1 points apart, under the 5 points the site counts as level.
| Opus 5 | Fable 5 | |||
|---|---|---|---|---|
| Low | Max | Low | Max | |
| Score | 87 | 91 | 91 | 94 |
| Web pages | 90 | 93 | 92 | 93 |
| Code | 100 | 100 | 100 | 100 |
| Planning | 65 | 76 | 77 | 94 |
| Everyday | 93 | 93 | 96 | 90 |
| Checks passed | 153 of 172 | 158 of 172 | 159 of 172 | 162 of 172 |
| Median time | 17 s | 57 s | 15 s | 1 min 15 s |
| Median cost | 4.7¢ | 13¢ | 7.8¢ | 33¢ |
Scored higher by 5 points or more: Fable 5 on planning and architecture, 94 to 76. Within 5 points on building web pages, code and everyday.
Opus 5’s median max effort answer took 57 s and cost 13¢ at API prices, and Fable 5’s median max effort answer took 1 min 15 s and cost 33¢. A median Fable 5 answer cost 2.6 times as much.
Of the 24 prompts at max effort, Opus 5 passed more checks on 4, Fable 5 on 7, and the two passed the same number on 13.
Where they differed
Prompts where the two passed a different number of checks at max effort, widest gap first.
A02 Should we switch to microservices?
Opus 5 5 of 8, Fable 5 8 of 8.
Passed by Fable 5, missed by Opus 5: Says up front: not now, keep the monolith; Stays under the 350-word limit; Suggests a middle path: clear modules inside the monolith
S02 A story of exactly 100 words
Opus 5 5 of 6, Fable 5 3 of 6.
Passed by Opus 5, missed by Fable 5: Exactly 100 words; Ends with the sentence 'The light stayed on.'; Gives only the story: no title, introduction or word-count note
Passed by Fable 5, missed by Opus 5: Uses the words 'kettle' and 'fog', and names the lighthouse or its keeper
B01 A landing page for a small invoicing app
Opus 5 9 of 9, Fable 5 7 of 9.
Passed by Opus 5, missed by Fable 5: One h1, no skipped heading levels, text 16 px or larger, enough contrast; The page looks finished: clear hierarchy, pricing easy to compare, nothing broken
B03 A kanban board that survives a reload
Opus 5 7 of 9, Fable 5 9 of 9.
Passed by Fable 5, missed by Opus 5: After a reload the board is exactly as it was; No sideways scroll on a 375 px phone, even with very long titles
H02 Summing up a book that doesn't exist
Opus 5 4 of 5, Fable 5 5 of 5.
Passed by Fable 5, missed by Opus 5: Under 200 words
R02 Did cycling in town really double?
Opus 5 5 of 6, Fable 5 4 of 6.
Passed by Opus 5, missed by Fable 5: Puts the town-wide rise at about 50%
5 more prompts differed. Each model page shows every answer.
Opus 5 and Haiku 4.5
Opus 5 scored 91 at max effort and Haiku 4.5 72 at its one setting, 18.9 points apart.
| Opus 5 | Haiku 4.5 | ||
|---|---|---|---|
| Low | Max | Only | |
| Score | 87 | 91 | 72 |
| Web pages | 90 | 93 | 63 |
| Code | 100 | 100 | 76 |
| Planning | 65 | 76 | 77 |
| Everyday | 93 | 93 | 71 |
| Checks passed | 153 of 172 | 158 of 172 | 124 of 172 |
| Median time | 17 s | 57 s | 6.8 s |
| Median cost | 4.7¢ | 13¢ | 0.5¢ |
Scored higher by 5 points or more: Opus 5 on building web pages, 93 to 63; Opus 5 on code, 100 to 76; Opus 5 on everyday, 93 to 71. Within 5 points on planning and architecture.
Opus 5’s median max effort answer took 57 s and cost 13¢ at API prices, and Haiku 4.5’s median answer took 6.8 s and cost 0.5¢. A median Opus 5 answer cost 27.9 times as much.
Of the 24 prompts, Opus 5 passed more checks on 12, Haiku 4.5 on 4, and the two passed the same number on 8.
Where they differed
Prompts where the two passed a different number of checks, each at its highest setting, widest gap first.
F05 SQL: which customers left and came back
Opus 5 7 of 7, Haiku 4.5 1 of 7.
Passed by Opus 5, missed by Haiku 4.5: It gets the example in the question right; Periods that touch or overlap are not counted as gaps; Several gaps per customer, same-name customers, sort order; A period inside a longer one does not hide or invent a gap; Open-ended (NULL end) periods never produce a gap after them; Days without access is exact across month, leap and year ends
I01 Check a market receipt adds up
Opus 5 5 of 5, Haiku 4.5 1 of 5.
Passed by Opus 5, missed by Haiku 4.5: Says the tomatoes line is wrong and should be 2.95; Gives the right total, 27.01; Says the overcharge is 1.00; Does not wrongly call any other line a mistake
F03 Some subscriptions renew a day early
Opus 5 7 of 7, Haiku 4.5 2 of 7.
Passed by Opus 5, missed by Haiku 4.5: Uses the customer's local date (Auckland, Kolkata, LA, London); Explains both causes: the UTC date and drifting month to month; Signups on the 29th to 31st keep their day after short months; Leap and non-leap Februaries are handled, including 2100; Signups just before or after a clock change get the right local date
B02 Build a dashboard from a picture of its design
Opus 5 7 of 7, Haiku 4.5 3 of 7.
Passed by Opus 5, missed by Haiku 4.5: Sidebar, active item, accent blue and the three status pills match the picture; Chart bars run Mon to Sun with heights in proportion to their values; On a phone: no sideways scroll, cards stacked, sidebar hidden until a menu button opens it; Side by side with the picture, a designer would accept it as a faithful build
B04 Minesweeper in one HTML file
Opus 5 8 of 9, Haiku 4.5 3 of 9.
Passed by Opus 5, missed by Haiku 4.5: Right click flags and unflags, the counter follows, a flagged square stays shut; Hitting a mine says Game over, shows all 10 mines as 💣 and stops the board; Opening every safe square shows You win! (and not before); New game resets the board, counter and message; Fits a 375 px phone with squares of at least 32 px, no sideways scroll
K02 Answer questions from a bread maker manual
Opus 5 6 of 6, Haiku 4.5 3 of 6.
Passed by Opus 5, missed by Haiku 4.5: Question 1: gets 4 hours 15 minutes (the footnote's extra 15 minutes included); Question 2: notices the delay needed is over the 13-hour maximum; Question 5: says the manual doesn't say how long keep-warm lasts (no made-up number)
10 more prompts differed. Each model page shows every answer.
Sonnet 5 and Fable 5
At max effort, Sonnet 5 scored 83 and Fable 5 94, 10.9 points apart. At low effort, Sonnet 5 scored 88 and Fable 5 91, 2.9 points apart, under the 5 points the site counts as level.
| Sonnet 5 | Fable 5 | |||
|---|---|---|---|---|
| Low | Max | Low | Max | |
| Score | 88 | 83 | 91 | 94 |
| Web pages | 79 | 74 | 92 | 93 |
| Code | 89 | 100 | 100 | 100 |
| Planning | 94 | 63 | 77 | 94 |
| Everyday | 92 | 96 | 96 | 90 |
| Checks passed | 151 of 172 | 146 of 172 | 159 of 172 | 162 of 172 |
| Median time | 15 s | 4 min 0 s | 15 s | 1 min 15 s |
| Median cost | 2.1¢ | 24¢ | 7.8¢ | 33¢ |
Scored higher by 5 points or more: Fable 5 on building web pages, 93 to 74; Fable 5 on planning and architecture, 94 to 63; Sonnet 5 on everyday, 96 to 90. Within 5 points on code.
Sonnet 5’s median max effort answer took 4 min 0 s and cost 24¢ at API prices, and Fable 5’s median max effort answer took 1 min 15 s and cost 33¢. A median Fable 5 answer cost 1.4 times as much.
Of the 24 prompts at max effort, Sonnet 5 passed more checks on 2, Fable 5 on 5, and the two passed the same number on 17.
Where they differed
Prompts where the two passed a different number of checks at max effort, widest gap first.
B03 A kanban board that survives a reload
Sonnet 5 0 of 9, Fable 5 9 of 9.
Passed by Fable 5, missed by Sonnet 5: The page loads with no errors and has the three named columns; Adding a card works with Enter and with the Add card button; Empty or space-only titles are ignored; The counts update, and Edit and Delete work; After a reload the board is exactly as it was; Dragging moves cards between columns and reorders them, and it survives a reload; A card can be moved with the arrow keys and keeps focus; Titles like <b>hi</b> & <img ...> show exactly as typed and run nothing; No sideways scroll on a 375 px phone, even with very long titles
A03 Rename a column with no downtime
Sonnet 5 0 of 9, Fable 5 8 of 9.
Passed by Fable 5, missed by Sonnet 5: Gives numbered steps; Adds a new handle column; Keeps both columns in step with a trigger or dual writes; Copies the existing 40 million rows in batches; Builds the unique index with CREATE INDEX CONCURRENTLY; Does not use RENAME COLUMN as the way to move while old code needs username; Keeps returning username to old mobile apps until they are gone; Gives a way back at every step
S02 A story of exactly 100 words
Sonnet 5 6 of 6, Fable 5 3 of 6.
Passed by Sonnet 5, missed by Fable 5: Exactly 100 words; Ends with the sentence 'The light stayed on.'; Gives only the story: no title, introduction or word-count note
A02 Should we switch to microservices?
Sonnet 5 6 of 8, Fable 5 8 of 8.
Passed by Fable 5, missed by Sonnet 5: Says when splitting would make sense; Suggests a middle path: clear modules inside the monolith
B01 A landing page for a small invoicing app
Sonnet 5 9 of 9, Fable 5 7 of 9.
Passed by Sonnet 5, missed by Fable 5: One h1, no skipped heading levels, text 16 px or larger, enough contrast; The page looks finished: clear hierarchy, pricing easy to compare, nothing broken
B05 Fix three layout bugs on a garden page
Sonnet 5 6 of 7, Fable 5 7 of 7.
Passed by Fable 5, missed by Sonnet 5: On a 375 px phone the page no longer scrolls sideways
1 more prompt differed. Each model page shows every answer.
Sonnet 5 and Haiku 4.5
Sonnet 5 scored 83 at max effort and Haiku 4.5 72 at its one setting, 11.4 points apart.
| Sonnet 5 | Haiku 4.5 | ||
|---|---|---|---|
| Low | Max | Only | |
| Score | 88 | 83 | 72 |
| Web pages | 79 | 74 | 63 |
| Code | 89 | 100 | 76 |
| Planning | 94 | 63 | 77 |
| Everyday | 92 | 96 | 71 |
| Checks passed | 151 of 172 | 146 of 172 | 124 of 172 |
| Median time | 15 s | 4 min 0 s | 6.8 s |
| Median cost | 2.1¢ | 24¢ | 0.5¢ |
Scored higher by 5 points or more: Sonnet 5 on building web pages, 74 to 63; Sonnet 5 on code, 100 to 76; Haiku 4.5 on planning and architecture, 77 to 63; Sonnet 5 on everyday, 96 to 71.
Sonnet 5’s median max effort answer took 4 min 0 s and cost 24¢ at API prices, and Haiku 4.5’s median answer took 6.8 s and cost 0.5¢. A median Sonnet 5 answer cost 53.2 times as much.
Of the 24 prompts, Sonnet 5 passed more checks on 12, Haiku 4.5 on 4, and the two passed the same number on 8.
Where they differed
Prompts where the two passed a different number of checks, each at its highest setting, widest gap first.
B03 A kanban board that survives a reload
Sonnet 5 0 of 9, Haiku 4.5 8 of 9.
Passed by Haiku 4.5, missed by Sonnet 5: The page loads with no errors and has the three named columns; Adding a card works with Enter and with the Add card button; Empty or space-only titles are ignored; The counts update, and Edit and Delete work; After a reload the board is exactly as it was; A card can be moved with the arrow keys and keeps focus; Titles like <b>hi</b> & <img ...> show exactly as typed and run nothing; No sideways scroll on a 375 px phone, even with very long titles
F05 SQL: which customers left and came back
Sonnet 5 7 of 7, Haiku 4.5 1 of 7.
Passed by Sonnet 5, missed by Haiku 4.5: It gets the example in the question right; Periods that touch or overlap are not counted as gaps; Several gaps per customer, same-name customers, sort order; A period inside a longer one does not hide or invent a gap; Open-ended (NULL end) periods never produce a gap after them; Days without access is exact across month, leap and year ends
I01 Check a market receipt adds up
Sonnet 5 5 of 5, Haiku 4.5 1 of 5.
Passed by Sonnet 5, missed by Haiku 4.5: Says the tomatoes line is wrong and should be 2.95; Gives the right total, 27.01; Says the overcharge is 1.00; Does not wrongly call any other line a mistake
A03 Rename a column with no downtime
Sonnet 5 0 of 9, Haiku 4.5 7 of 9.
Passed by Haiku 4.5, missed by Sonnet 5: Gives numbered steps; Adds a new handle column; Keeps both columns in step with a trigger or dual writes; Copies the existing 40 million rows in batches; Builds the unique index with CREATE INDEX CONCURRENTLY; Does not use RENAME COLUMN as the way to move while old code needs username; Drops the old column last, after a waiting period
F03 Some subscriptions renew a day early
Sonnet 5 7 of 7, Haiku 4.5 2 of 7.
Passed by Sonnet 5, missed by Haiku 4.5: Uses the customer's local date (Auckland, Kolkata, LA, London); Explains both causes: the UTC date and drifting month to month; Signups on the 29th to 31st keep their day after short months; Leap and non-leap Februaries are handled, including 2100; Signups just before or after a clock change get the right local date
B04 Minesweeper in one HTML file
Sonnet 5 9 of 9, Haiku 4.5 3 of 9.
Passed by Sonnet 5, missed by Haiku 4.5: Numbers are right and an empty square opens its neighbours; Right click flags and unflags, the counter follows, a flagged square stays shut; Hitting a mine says Game over, shows all 10 mines as 💣 and stops the board; Opening every safe square shows You win! (and not before); New game resets the board, counter and message; Fits a 375 px phone with squares of at least 32 px, no sideways scroll
10 more prompts differed. Each model page shows every answer.
Fable 5 and Haiku 4.5
Fable 5 scored 94 at max effort and Haiku 4.5 72 at its one setting, 22.3 points apart.
| Fable 5 | Haiku 4.5 | ||
|---|---|---|---|
| Low | Max | Only | |
| Score | 91 | 94 | 72 |
| Web pages | 92 | 93 | 63 |
| Code | 100 | 100 | 76 |
| Planning | 77 | 94 | 77 |
| Everyday | 96 | 90 | 71 |
| Checks passed | 159 of 172 | 162 of 172 | 124 of 172 |
| Median time | 15 s | 1 min 15 s | 6.8 s |
| Median cost | 7.8¢ | 33¢ | 0.5¢ |
Scored higher by 5 points or more: Fable 5 on building web pages, 93 to 63; Fable 5 on code, 100 to 76; Fable 5 on planning and architecture, 94 to 77; Fable 5 on everyday, 90 to 71.
Fable 5’s median max effort answer took 1 min 15 s and cost 33¢ at API prices, and Haiku 4.5’s median answer took 6.8 s and cost 0.5¢. A median Fable 5 answer cost 72.6 times as much.
Of the 24 prompts, Fable 5 passed more checks on 15, Haiku 4.5 on 2, and the two passed the same number on 7.
Where they differed
Prompts where the two passed a different number of checks, each at its highest setting, widest gap first.
F05 SQL: which customers left and came back
Fable 5 7 of 7, Haiku 4.5 1 of 7.
Passed by Fable 5, missed by Haiku 4.5: It gets the example in the question right; Periods that touch or overlap are not counted as gaps; Several gaps per customer, same-name customers, sort order; A period inside a longer one does not hide or invent a gap; Open-ended (NULL end) periods never produce a gap after them; Days without access is exact across month, leap and year ends
I01 Check a market receipt adds up
Fable 5 5 of 5, Haiku 4.5 1 of 5.
Passed by Fable 5, missed by Haiku 4.5: Says the tomatoes line is wrong and should be 2.95; Gives the right total, 27.01; Says the overcharge is 1.00; Does not wrongly call any other line a mistake
F03 Some subscriptions renew a day early
Fable 5 7 of 7, Haiku 4.5 2 of 7.
Passed by Fable 5, missed by Haiku 4.5: Uses the customer's local date (Auckland, Kolkata, LA, London); Explains both causes: the UTC date and drifting month to month; Signups on the 29th to 31st keep their day after short months; Leap and non-leap Februaries are handled, including 2100; Signups just before or after a clock change get the right local date
B04 Minesweeper in one HTML file
Fable 5 9 of 9, Haiku 4.5 3 of 9.
Passed by Fable 5, missed by Haiku 4.5: Numbers are right and an empty square opens its neighbours; Right click flags and unflags, the counter follows, a flagged square stays shut; Hitting a mine says Game over, shows all 10 mines as 💣 and stops the board; Opening every safe square shows You win! (and not before); New game resets the board, counter and message; Fits a 375 px phone with squares of at least 32 px, no sideways scroll
K02 Answer questions from a bread maker manual
Fable 5 6 of 6, Haiku 4.5 3 of 6.
Passed by Fable 5, missed by Haiku 4.5: Question 1: gets 4 hours 15 minutes (the footnote's extra 15 minutes included); Question 2: notices the delay needed is over the 13-hour maximum; Question 5: says the manual doesn't say how long keep-warm lasts (no made-up number)
P02 Filling a week of volunteer shifts
Fable 5 6 of 6, Haiku 4.5 3 of 6.
Passed by Fable 5, missed by Haiku 4.5: Nobody is on a shift they can't do; Everyone works 3 or 4 shifts, and Osei exactly 2; Wren and Tomasz are never together, and Lindiwe is always with Priya
11 more prompts differed. Each model page shows every answer.
Written by an AI model from these results; every figure in these sentences is filled in by a program. How it was made.
Comparisons with a page of their own
- Sonnet 5.5 vs Opus 5.5At max effort, Sonnet 5.5 scored 92 and Opus 5.5 93, 1.2 points apart, under the 5 points the site counts as level.
- Sonnet 5.5 vs Sonnet 5At max effort, Sonnet 5.5 scored 92 and Sonnet 5 83, 8.9 points apart.
- Sonnet 5.5 vs Haiku 4.5Sonnet 5.5 scored 92 at max effort and Haiku 4.5 72 at its one setting, 20.3 points apart.
- Opus 5.5 vs Fable 5.1At max effort, Opus 5.5 scored 93 and Fable 5.1 97, 3.6 points apart, under the 5 points the site counts as level.
- Opus 5.5 vs Opus 5At max effort, Opus 5.5 scored 93 and Opus 5 91, 2.6 points apart, under the 5 points the site counts as level.