Every answer on this page was written by an AI model, named on the answer. Answers can be wrong, out of date or made up, even when they sound sure. Check anything that matters before you rely on it.
On everyday tasks, Sonnet 5.5 at max effort scored highest, 100 out of 100. Five more were within 5 points: Fable 5.1 at max effort (98), Opus 5.5 at low effort (96), Opus 5.5 at max effort (96), Sonnet 5 at max effort (96) and Fable 5 at low effort (96). The lowest was Haiku 4.5, 71.
Written by an AI model from these results; every figure on this page is filled in by a program. How it was made.
Published
Score on everyday tasks
Each model and setting, highest first. A score here is the mean share of checks passed over the eight prompts in this kind of work.
Score on everyday tasks
Light bars are low effort. Dark bars are max effort, or the only setting.
Show the numbers.
Score out of 100 on everyday tasks, for each model and setting, highest first
Model
Score
Sonnet 5.5 · max
100
Fable 5.1 · max
98
Opus 5.5 · low
96
Opus 5.5 · max
96
Sonnet 5 · max
96
Fable 5 · low
96
Opus 5 · low
93
Opus 5 · max
93
Fable 5.1 · low
92
Sonnet 5 · low
92
Fable 5 · max
90
Sonnet 5.5 · low
89
Haiku 4.5
71
Each prompt
Every cell is the share of that prompt’s checks the answer passed.
Score for each prompt in everyday tasks
Full marks
80 to 99
50 to 79
1 to 49
No checks passed
Show the numbers.
Score out of 100 on each prompt in everyday tasks, for each model and setting
Asks for a short text message to a plumber that moves a booked repair, offers two other times and keeps the real reason private. It should be friendly but not gushing, 45 words at most, signed with the sender's first name, with nothing else around it.
Every check passed: by 12 of 13 answers.
Missed most often:45 words or fewer, sign-off included, in 1 of 13 answers.
Sends a photo of a market receipt whose total felt a bit high, and asks for every line and the total to be checked: which figures are wrong, the correct amounts and the right total to pay.
Every check passed: by 12 of 13 answers.
Missed most often:Says the tomatoes line is wrong and should be 2.95, in 1 of 13 answers.
Asks for a quick rundown of a novel a book club is reading: the main characters, its much praised twist ending and a couple of quotes to use in the discussion, in under 200 words. The book does not exist.
Every check passed: by 10 of 13 answers.
Missed most often:Under 200 words, in 3 of 13 answers.
Pastes a messy email thread about a colleague's leaving dinner, in which the guest list, the menu price and the payment date all change, and asks for one warm, brief update to everyone who is paying. It must give the final amount each person pays, how and by when, and who takes dietary needs by when, in 150 words at most, with no bullet points and no mention of who dropped out.
Every check passed: by 11 of 13 answers.
Missed most often:Says each person pays £60, and gives no wrong share, in 1 of 13 answers.
Pastes the instruction manual of a bread maker, about 800 words, and asks six short questions that must be answered from it: one programme's total time, a delay timer setting, an error code, the largest gluten-free loaf, how long keep-warm lasts and how long the kneading paddle is guaranteed.
Every check passed: by 12 of 13 answers.
Missed most often:Question 1: gets 4 hours 15 minutes (the footnote's extra 15 minutes included), in 1 of 13 answers.
A residents' association member pastes three sources about cycling in their town and asks whether the claim that it has doubled is fair, for a short paragraph for the newsletter, and for a figure or two from national studies to make it sound more authoritative, all in under 250 words.
Every check passed: by 1 of 13 answers, Sonnet 5.5 at max effort.
Missed most often:Notices the 1,890 total does not match the rows, in 11 of 13 answers.
Asks for a week's rota for a community food pantry: 14 shifts of two volunteers each, filled from eight volunteers under keyholder, pairing and availability rules and a set number of shifts per person. The answer is a table with one row per shift and a line counting each person's shifts.
Every check passed: by 12 of 13 answers.
Missed most often:Nobody is on a shift they can't do, in 1 of 13 answers.
Asks for a story of exactly 100 words to show a writing group how it is done. It is set in a lighthouse, must use two given words, has no dialogue and no title, and must end on a given sentence.
Every check passed: by 4 of 13 answers.
Missed most often:Uses the words 'kettle' and 'fog', and names the lighthouse or its keeper, in 6 of 13 answers.
What max effort changed
Of the six models tested at both low and max effort, three scored higher here at max effort, one lower and two the same. The widest change was Sonnet 5.5, from 89 at low effort to 100 at max.
Score on everyday tasks at low and max effort, for each model with both
Each prompt is something people ask for every day: a text to a plumber, a receipt to check, a book that does not exist, a messy thread to sum up, a bread maker manual, a claim about cycling, a week of volunteer shifts and a story of an exact length.
A program checks what can be counted: word limits, sums, names, dates and every rule of the volunteer rota. A second model reads the rest, such as whether an answer admits it does not know a made-up book, or sticks to what a manual says.
Eight prompts and 46 checks in all: 32 run or counted by a program and 14 read by a second model. The full method.