Every answer on this page was written by an AI model, named on the answer. Answers can be wrong, out of date or made up, even when they sound sure. Check anything that matters before you rely on it.
Max effort results are being collected again. We capped every answer at 32,000 output tokens, below the output limit of six of the seven models, so many max effort answers ran out of room while still thinking and scored zero. Low effort results and Haiku 4.5 are not affected. Ranks are hidden until the new answers are in.
Analytics
7 models, 24 prompts, one answer to each per setting. A score is the share of checks an answer passed.
Models tested
7
Haiku 4.5, Fable 5, Sonnet 5, Opus 5, Fable 5.1, Opus 5.5, Sonnet 5.5
Prompts
24
In 4 kinds of work. The list does not change.
Checks marked
2,236
Every check on every answer, 1,784 passed.
Answers collected
312
One per prompt for each model and setting.
Checks passed
80%
Of all checks marked, across every answer.
Time to answer everything
7.5 hours
Median 35 s for one answer.
Output tokens
2.7M
2,744,983 in all, 83% of them thinking.
Cost at API prices
$67.66
For every answer, at the prices of 30 Sep 2026. Not what anyone paid.
Score by model
Each model at its highest effort setting. The ring is the same model at low effort. Ranks are hidden while the max effort answers are collected again.
Score by model
Ring: low effort. Dot: max effort, or the only setting.
Show the numbers.
Score out of 100 for each model. Where a model has an effort setting, the ring is low effort and the dot is max effort.
Model
Low effort
Max effort or only setting
Fable 5
91
95
Opus 5
87
91
Haiku 4.5
No effort setting
72
Fable 5.1
95
72
Sonnet 5
88
54
Sonnet 5.5
93
49
Opus 5.5
90
49
Score by kind of work
The four kinds of work count equally in the overall score, so a strong showing in one cannot hide a weak one in another.
Score by kind of work
Building web pages
Best: Fable 5.1 · low, 98, and 2 more within 5 points.
Code
Best: Fable 5 · low, 100, and 6 more within 5 points.
Planning and architecture
Best: Sonnet 5 · low, 94, and 3 more within 5 points.
Everyday
Best: Fable 5.1 · max, 100, and 5 more within 5 points.
Light bars are low effort. Dark bars are max effort, or the only setting.
Show the numbers.
Score out of 100 for each model and setting, by kind of work
Model
Building web pages
Code
Planning and architecture
Everyday
Haiku 4.5
63
76
77
71
Fable 5 · low
92
100
77
96
Fable 5 · max
95
100
91
94
Sonnet 5 · low
79
89
94
92
Sonnet 5 · max
20
57
57
81
Opus 5 · low
90
100
65
93
Opus 5 · max
76
100
91
96
Fable 5.1 · low
98
100
91
92
Fable 5.1 · max
40
86
60
100
Opus 5.5 · low
90
98
76
96
Opus 5.5 · max
0
71
25
98
Sonnet 5.5 · low
98
98
85
89
Sonnet 5.5 · max
0
71
25
100
Score for each prompt
Every cell is the share of that prompt's checks the answer passed. The number is always printed; the shade only follows it. Prompts are in the four kinds of work.
Score for each prompt
Full marks
80 to 99
50 to 79
1 to 49
No checks passed
Show the numbers.
Score out of 100 for each prompt, for each model and setting
Prompt
Haiku 4.5
Fable 5 low
Fable 5 max
Sonnet 5 low
Sonnet 5 max
Opus 5 low
Opus 5 max
Fable 5.1 low
Fable 5.1 max
Opus 5.5 low
Opus 5.5 max
Sonnet 5.5 low
Sonnet 5.5 max
B01 A landing page for a small invoicing app
67
100
100
78
0
100
89
89
100
100
0
100
0
B02 Build a dashboard from a picture of its design
43
71
86
86
0
86
100
100
0
86
0
100
0
B03 A kanban board that survives a reload
89
100
89
89
0
100
0
100
0
78
0
89
0
B04 Minesweeper in one HTML file
33
89
100
56
100
67
89
100
0
89
0
100
0
B05 Fix three layout bugs on a garden page
86
100
100
86
0
100
100
100
100
100
0
100
0
F01 Load more shows some posts twice and skips others
100
100
100
100
100
100
100
100
100
100
100
100
100
F02 The search box shows results for an old query
100
100
100
100
100
100
100
100
100
100
100
100
100
F03 Some subscriptions renew a day early
29
100
100
86
100
100
100
100
100
100
100
100
100
F04 Review a pull request before merging
86
100
100
100
100
100
100
100
100
86
100
86
100
F05 SQL: which customers left and came back
14
100
100
57
0
100
100
100
0
100
100
100
100
F06 A contacts import that survives real CSV files
100
100
100
78
0
100
100
100
100
100
0
100
0
F07 Make a slow script fast
100
100
100
100
0
100
100
100
100
100
0
100
0
A01 Plan a booking app for a dog groomer
67
100
89
89
89
89
89
100
0
78
0
78
0
A02 Should we switch to microservices?
75
88
88
88
88
75
100
88
88
88
100
88
100
A03 Rename a column with no downtime
78
44
89
100
0
44
89
89
67
78
0
89
0
A04 Turn a messy feature brief into a spec
88
75
100
100
50
50
88
88
88
63
0
88
0
W01 Text the plumber to reschedule
100
100
100
100
100
100
100
100
100
100
100
83
100
I01 Check a market receipt adds up
20
100
100
100
100
100
100
100
100
100
100
100
100
H02 Summing up a book that doesn't exist
100
100
100
100
80
80
100
100
100
100
100
80
100
W02 Turn a messy thread into one clear update
83
100
100
100
100
100
100
83
100
100
100
100
100
K02 Answer questions from a bread maker manual
50
100
100
100
100
100
100
100
100
100
100
100
100
R02 Did cycling in town really double?
83
83
50
50
83
83
83
50
100
83
83
67
100
P02 Filling a week of volunteer shifts
50
100
100
100
0
100
100
100
100
100
100
100
100
S02 A story of exactly 100 words
83
83
100
83
83
83
83
100
100
83
100
83
100
The three layers
Floor checks are the ones nearly every answer should pass, middle checks reward care, and top checks separate the strong answers. How they are marked.
Haiku 4.5
Floor47/49, 96% passed
Middle45/58, 78% passed
Top32/65, 49% passed
Fable 5 · low
Floor47/49, 96% passed
Middle54/58, 93% passed
Top58/65, 89% passed
Fable 5 · max
Floor48/49, 98% passed
Middle56/58, 97% passed
Top60/65, 92% passed
Sonnet 5 · low
Floor46/49, 94% passed
Middle51/58, 88% passed
Top54/65, 83% passed
Sonnet 5 · max
Floor29/49, 59% passed
Middle30/58, 52% passed
Top34/65, 52% passed
Opus 5 · low
Floor45/49, 92% passed
Middle51/58, 88% passed
Top57/65, 88% passed
Opus 5 · max
Floor45/49, 92% passed
Middle52/58, 90% passed
Top59/65, 91% passed
Fable 5.1 · low
Floor47/49, 96% passed
Middle55/58, 95% passed
Top62/65, 95% passed
Fable 5.1 · max
Floor39/49, 80% passed
Middle42/58, 72% passed
Top45/65, 69% passed
Opus 5.5 · low
Floor46/49, 94% passed
Middle51/58, 88% passed
Top60/65, 92% passed
Opus 5.5 · max
Floor28/49, 57% passed
Middle28/58, 48% passed
Top32/65, 49% passed
Sonnet 5.5 · low
Floor45/49, 92% passed
Middle53/58, 91% passed
Top62/65, 95% passed
Sonnet 5.5 · max
Floor29/49, 59% passed
Middle28/58, 48% passed
Top32/65, 49% passed
Across all answers, 85% of floor checks passed, 79% of middle checks and 77% of top checks. Fable 5.1 · low and Sonnet 5.5 · low passed the most top checks, 62 of 65.
Show the numbers.
Checks passed in each layer, for each model and setting
Model and setting
Floor
Middle
Top
Haiku 4.5
47 of 49
45 of 58
32 of 65
Fable 5 · low
47 of 49
54 of 58
58 of 65
Fable 5 · max
48 of 49
56 of 58
60 of 65
Sonnet 5 · low
46 of 49
51 of 58
54 of 65
Sonnet 5 · max
29 of 49
30 of 58
34 of 65
Opus 5 · low
45 of 49
51 of 58
57 of 65
Opus 5 · max
45 of 49
52 of 58
59 of 65
Fable 5.1 · low
47 of 49
55 of 58
62 of 65
Fable 5.1 · max
39 of 49
42 of 58
45 of 65
Opus 5.5 · low
46 of 49
51 of 58
60 of 65
Opus 5.5 · max
28 of 49
28 of 58
32 of 65
Sonnet 5.5 · low
45 of 49
53 of 58
62 of 65
Sonnet 5.5 · max
29 of 49
28 of 58
32 of 65
All answers
541 of 637, 85%
596 of 754, 79%
647 of 845, 77%
What max effort changed
Only models tested at both low and max effort appear here. Across these 6 models, max effort scored higher on 26 prompts, lower on 42 and the same on 76.
Fable 5
91to95 points, low to max effort
Max passed 5 more checks, cost 3.8 times as much and took 4.5 times as long.
Max scored higher
5
Max scored lower
3
No difference
16
Cost
3.8 times
Time
4.5 times
Fable 5: change by prompt
Points gained or lost at max effort, only for prompts that scored differently.
Show the numbers.
Change in score for each prompt that scored differently, Fable 5, low to max effort
Prompt
Low effort
Max effort
Change, points
A03 Rename a column with no downtime
44
89
+44
A04 Turn a messy feature brief into a spec
75
100
+25
S02 A story of exactly 100 words
83
100
+17
B02 Build a dashboard from a picture of its design
71
86
+14
B04 Minesweeper in one HTML file
89
100
+11
B03 A kanban board that survives a reload
100
89
-11
A01 Plan a booking app for a dog groomer
100
89
-11
R02 Did cycling in town really double?
83
50
-33
Sonnet 5
88to54 points, low to max effort
Max passed 58 fewer checks, cost 6.2 times as much and took 7.8 times as long.
Max scored higher
3
Max scored lower
11
No difference
10
Cost
6.2 times
Time
7.8 times
Sonnet 5: change by prompt
Points gained or lost at max effort, only for prompts that scored differently.
Show the numbers.
Change in score for each prompt that scored differently, Sonnet 5, low to max effort
Prompt
Low effort
Max effort
Change, points
B04 Minesweeper in one HTML file
56
100
+44
R02 Did cycling in town really double?
50
83
+33
F03 Some subscriptions renew a day early
86
100
+14
H02 Summing up a book that doesn't exist
100
80
-20
A04 Turn a messy feature brief into a spec
100
50
-50
F05 SQL: which customers left and came back
57
0
-57
B01 A landing page for a small invoicing app
78
0
-78
F06 A contacts import that survives real CSV files
78
0
-78
B02 Build a dashboard from a picture of its design
86
0
-86
B05 Fix three layout bugs on a garden page
86
0
-86
B03 A kanban board that survives a reload
89
0
-89
F07 Make a slow script fast
100
0
-100
A03 Rename a column with no downtime
100
0
-100
P02 Filling a week of volunteer shifts
100
0
-100
Opus 5
87to91 points, low to max effort
Max passed 3 more checks, cost 3.4 times as much and took 3.8 times as long.
Max scored higher
6
Max scored lower
2
No difference
16
Cost
3.4 times
Time
3.8 times
Opus 5: change by prompt
Points gained or lost at max effort, only for prompts that scored differently.
Show the numbers.
Change in score for each prompt that scored differently, Opus 5, low to max effort
Prompt
Low effort
Max effort
Change, points
A03 Rename a column with no downtime
44
89
+44
A04 Turn a messy feature brief into a spec
50
88
+38
A02 Should we switch to microservices?
75
100
+25
B04 Minesweeper in one HTML file
67
89
+22
H02 Summing up a book that doesn't exist
80
100
+20
B02 Build a dashboard from a picture of its design
86
100
+14
B01 A landing page for a small invoicing app
100
89
-11
B03 A kanban board that survives a reload
100
0
-100
Fable 5.1
95to72 points, low to max effort
Max passed 38 fewer checks, cost 5.9 times as much and took 6.5 times as long.
Max scored higher
3
Max scored lower
6
No difference
15
Cost
5.9 times
Time
6.5 times
Fable 5.1: change by prompt
Points gained or lost at max effort, only for prompts that scored differently.
Show the numbers.
Change in score for each prompt that scored differently, Fable 5.1, low to max effort
Prompt
Low effort
Max effort
Change, points
R02 Did cycling in town really double?
50
100
+50
W02 Turn a messy thread into one clear update
83
100
+17
B01 A landing page for a small invoicing app
89
100
+11
A03 Rename a column with no downtime
89
67
-22
B02 Build a dashboard from a picture of its design
100
0
-100
B03 A kanban board that survives a reload
100
0
-100
B04 Minesweeper in one HTML file
100
0
-100
F05 SQL: which customers left and came back
100
0
-100
A01 Plan a booking app for a dog groomer
100
0
-100
Opus 5.5
90to49 points, low to max effort
Max passed 69 fewer checks, cost 8.1 times as much and took 9.7 times as long.
Max scored higher
3
Max scored lower
10
No difference
11
Cost
8.1 times
Time
9.7 times
Opus 5.5: change by prompt
Points gained or lost at max effort, only for prompts that scored differently.
Show the numbers.
Change in score for each prompt that scored differently, Opus 5.5, low to max effort
Prompt
Low effort
Max effort
Change, points
S02 A story of exactly 100 words
83
100
+17
F04 Review a pull request before merging
86
100
+14
A02 Should we switch to microservices?
88
100
+13
A04 Turn a messy feature brief into a spec
63
0
-62
B03 A kanban board that survives a reload
78
0
-78
A01 Plan a booking app for a dog groomer
78
0
-78
A03 Rename a column with no downtime
78
0
-78
B02 Build a dashboard from a picture of its design
86
0
-86
B04 Minesweeper in one HTML file
89
0
-89
B01 A landing page for a small invoicing app
100
0
-100
B05 Fix three layout bugs on a garden page
100
0
-100
F06 A contacts import that survives real CSV files
100
0
-100
F07 Make a slow script fast
100
0
-100
Sonnet 5.5
93to49 points, low to max effort
Max passed 71 fewer checks, cost 9.4 times as much and took 11.3 times as long.
Max scored higher
6
Max scored lower
10
No difference
8
Cost
9.4 times
Time
11.3 times
Sonnet 5.5: change by prompt
Points gained or lost at max effort, only for prompts that scored differently.
Show the numbers.
Change in score for each prompt that scored differently, Sonnet 5.5, low to max effort
Prompt
Low effort
Max effort
Change, points
R02 Did cycling in town really double?
67
100
+33
H02 Summing up a book that doesn't exist
80
100
+20
W01 Text the plumber to reschedule
83
100
+17
S02 A story of exactly 100 words
83
100
+17
F04 Review a pull request before merging
86
100
+14
A02 Should we switch to microservices?
88
100
+13
A01 Plan a booking app for a dog groomer
78
0
-78
A04 Turn a messy feature brief into a spec
88
0
-87
B03 A kanban board that survives a reload
89
0
-89
A03 Rename a column with no downtime
89
0
-89
B01 A landing page for a small invoicing app
100
0
-100
B02 Build a dashboard from a picture of its design
100
0
-100
B04 Minesweeper in one HTML file
100
0
-100
B05 Fix three layout bugs on a garden page
100
0
-100
F06 A contacts import that survives real CSV files
100
0
-100
F07 Make a slow script fast
100
0
-100
Cost against score
Median cost of one answer at API prices of , against score. It is not what anyone paid. Up and to the left is more score for less.
Cost against score
Ring: low effort. Dot: max effort, or the only setting.
Show the numbers.
Median cost of one answer at API prices, and score, for each model and setting
Model
Median cost of one answer
Score
Haiku 4.5
0.5¢
72
Fable 5 · low
7.8¢
91
Fable 5 · max
29¢
95
Sonnet 5 · low
2.1¢
88
Sonnet 5 · max
24¢
54
Opus 5 · low
4.7¢
87
Opus 5 · max
14¢
91
Fable 5.1 · low
11¢
95
Fable 5.1 · max
62¢
72
Opus 5.5 · low
4.1¢
90
Opus 5.5 · max
55¢
49
Sonnet 5.5 · low
1.7¢
93
Sonnet 5.5 · max
30¢
49
Speed and tokens
Medians for one answer. Time and tokens never enter a score. Thinking tokens are counted inside the output tokens.
Time to answer
Seconds. Light bars are low effort. Dark bars are max effort, or the only setting.
Show the numbers.
Median time in seconds to the finished answer and to the first words, for each model and setting
Model
Time to the finished answer
Time to the first words
Haiku 4.5
6.8 s
1.4 s
Fable 5 · low
15 s
3.2 s
Fable 5 · max
1 min 9 s
49 s
Sonnet 5 · low
15 s
2.1 s
Sonnet 5 · max
3 min 26 s
not recorded
Opus 5 · low
17 s
4.2 s
Opus 5 · max
1 min 1 s
40 s
Fable 5.1 · low
20 s
6.3 s
Fable 5.1 · max
2 min 16 s
not recorded
Opus 5.5 · low
16 s
6.1 s
Opus 5.5 · max
4 min 0 s
not recorded
Sonnet 5.5 · low
8.8 s
1.9 s
Sonnet 5.5 · max
3 min 19 s
not recorded
Output tokens
Each bar is all the output tokens. The dark part is thinking, and the label ends with its share.
Show the numbers.
Output tokens for all answers, and the share of them that was thinking, for each model and setting
Model
Output tokens
Haiku 4.5
29k · 0%
Fable 5 · low
45k · 28%
Fable 5 · max
197k · 77%
Sonnet 5 · low
70k · 38%
Sonnet 5 · max
484k · 96%
Opus 5 · low
56k · 15%
Opus 5 · max
214k · 72%
Fable 5.1 · low
60k · 24%
Fable 5.1 · max
388k · 91%
Opus 5.5 · low
58k · 20%
Opus 5.5 · max
536k · 98%
Sonnet 5.5 · low
51k · 19%
Sonnet 5.5 · max
556k · 98%
Hardest prompts and most missed checks
The prompts with the lowest mean score over every answer, and the checks that most answers missed. Each links to the prompt.