Every answer on this page was written by an AI model, named on the answer. Answers can be wrong, out of date or made up, even when they sound sure. Check anything that matters before you rely on it.
On code, nine of the 13 models and settings scored 100 out of 100: Sonnet 5.5 at max effort, Opus 5.5 at max effort, Fable 5.1 at low effort, Fable 5.1 at max effort and five more. Two more were within 5 points: Sonnet 5.5 at low effort (98) and Opus 5.5 at low effort (98). The lowest was Haiku 4.5, 76.
Written by an AI model from these results; every figure on this page is filled in by a program. How it was made.
Published
Score on code
Each model and setting, highest first. A score here is the mean share of checks passed over the seven prompts in this kind of work.
Score on code
Light bars are low effort. Dark bars are max effort, or the only setting.
Show the numbers.
Score out of 100 on code, for each model and setting, highest first
Model
Score
Sonnet 5.5 · max
100
Opus 5.5 · max
100
Fable 5.1 · low
100
Fable 5.1 · max
100
Opus 5 · low
100
Opus 5 · max
100
Sonnet 5 · max
100
Fable 5 · low
100
Fable 5 · max
100
Sonnet 5.5 · low
98
Opus 5.5 · low
98
Sonnet 5 · low
89
Haiku 4.5
76
Each prompt
Every cell is the share of that prompt’s checks the answer passed.
Score for each prompt in code
Full marks
80 to 99
50 to 79
1 to 49
No checks passed
Show the numbers.
Score out of 100 on each prompt in code, for each model and setting
Prompt
Sonnet 5.5 low
Sonnet 5.5 max
Opus 5.5 low
Opus 5.5 max
Fable 5.1 low
Fable 5.1 max
Opus 5 low
Opus 5 max
Sonnet 5 low
Sonnet 5 max
Fable 5 low
Fable 5 max
Haiku 4.5
F01 Load more shows some posts twice and skips others
100
100
100
100
100
100
100
100
100
100
100
100
100
F02 The search box shows results for an old query
100
100
100
100
100
100
100
100
100
100
100
100
100
F03 Some subscriptions renew a day early
100
100
100
100
100
100
100
100
86
100
100
100
29
F04 Review a pull request before merging
86
100
86
100
100
100
100
100
100
100
100
100
86
F05 SQL: which customers left and came back
100
100
100
100
100
100
100
100
57
100
100
100
14
F06 A contacts import that survives real CSV files
Pastes a Python function that pages through a community board's posts, newest first, and reports that some posts show up twice while others never appear, more often on busy days. Many posts share exactly the same timestamp. It asks for the cause and a fixed function with the same name, arguments and return shape, where posts with the same time come higher id first.
Pastes a TypeScript controller for a search-as-you-type box with two bugs: typing fast can leave results for an earlier query on screen, and the loading spinner disappears too early. The API client cannot cancel requests. It asks for the whole fixed file with the same exports, the spinner on exactly while the newest search is loading, and code that Node can run with its built-in TypeScript support.
Pastes a JavaScript function that lists a subscription's monthly renewal dates, with its rules in a comment, and reports two symptoms: customers far from UTC renew a day early, and a customer who signed up on the last day of January has been charged on the 28th of every month since. It asks for the function fixed for Node 22 with no libraries, and a short note on what was wrong.
Every check passed: by 11 of 13 answers.
Missed most often:Explains both causes: the UTC date and drifting month to month, in 2 of 13 answers.
Pastes a pull request of about 170 lines that adds invoice endpoints to an Express and TypeScript API, and asks for a review before merging that lists what needs fixing in order of importance plus other points worth noting, in under 400 words.
Every check passed: by 10 of 13 answers.
Missed most often:Stays under the 400-word limit, in 2 of 13 answers.
Gives the schema of a subscriptions database and asks for one SQLite query that lists each gap in a customer's access that ended with them returning, with the last day of access, the day it came back and the number of days without it. Plan changes can make periods touch, overlap or sit inside one another, and none of that counts as a gap. A small example with its expected output is included.
Every check passed: by 11 of 13 answers.
Missed most often:It gets the example in the question right, in 2 of 13 answers.
Asks for a Python function, standard library only, that imports contacts from CSV files exported by many different tools and returns each contact's name, email and phone. The files vary in header names, comma or semicolon separators, a byte order mark, quoted fields with line breaks, Windows line endings and older files in Windows-1252. Rows without an email are skipped, and a repeated email keeps its first row.
Every check passed: by 12 of 13 answers.
Missed most often:Line breaks inside quoted fields are kept as one \n, in 1 of 13 answers.
Pastes a Python function that matches bank payments to open invoices by customer, amount and a due date within 7 days, and says it is far too slow on about 200,000 of each. It asks for a replacement that runs in a couple of seconds, returns exactly the same results for every input and leaves the lists it is given unchanged.
Every check passed: by every answer.
What max effort changed
Of the six models tested at both low and max effort, three scored higher here at max effort, no lower and three the same. The widest change was Sonnet 5, from 89 at low effort to 100 at max.
Score on code at low and max effort, for each model with both
Each prompt is a bug or a job from ordinary working code: posts that repeat when more are loaded, a search box that shows results for an old query, renewals a day early, a pull request to review, a SQL query, a contacts import and a script that is too slow.
The code in each answer is run in a sandbox against hidden tests the model never saw. The few other checks are counted by a program or read by a second model, such as whether a review names the real risk.
Seven prompts and 51 checks in all: 48 run or counted by a program and 3 read by a second model. The full method.