ModelLineup

Every answer on this page was written by an AI model, named on the answer. Answers can be wrong, out of date or made up, even when they sound sure. Check anything that matters before you rely on it.

Max effort results for this model are being collected again. We capped every answer at 32,000 output tokens, below this model’s own output limit, so many max effort answers ran out of room while still thinking and scored zero. Low effort results are not affected. The new max effort answers will replace these.

Models Sonnet 5.5

Claude Sonnet 5.5

Released
Tested
Settings
Low effort and Max effort

Sonnet 5.5

24 prompts, every check counted

Low effort

93

  • Floor45/49, 92% passed
  • Middle53/58, 91% passed
  • Top62/65, 95% passed

Median answer 8.8 s, 1.7¢.

Max effort

49

  • Floor29/49, 59% passed
  • Middle28/58, 48% passed
  • Top32/65, 49% passed

Median answer 3 min 19 s, 30¢.

Max passed 71 fewer checks, cost 9.4 times as much and took 11.3 times as long. Costs are at Anthropic’s API prices on .

What the results show

Written by an AI model from these results; every figure in it is filled in by a program. How it was made.

Sonnet 5.5 answered all 24 prompts at both efforts on 2 October 2026, scoring 93 at low effort and 49 at max. At max effort 10 answers ran out of room, every web page among them, with the output limit spent mostly on thinking.

Best at

  • Web pages at low effort: the landing page, the dashboard from the mockup, the Minesweeper game and the garden page fixes passed every page test. B01 B02 B04 B05
  • Code at low effort: it passed every hidden test on the paging bug, the search box race, the early renewals, the customers who left and came back, the contacts import and the slow script. F01 F02 F03 F05 F06 F07
  • Careful reading at both efforts: it found the wrong line on the market receipt and kept its bread maker answers to what the manual says. I01 K02
  • The volunteer rota and the update from a messy thread met every check at both efforts. P02 W02

Stumbled on

  • At max effort every web page answer ran out of room before any text was written, so there was nothing for the page tests to load. B01 B02 B03 B04 B05
  • Also at max effort the contacts import, the slow script, the booking plan, the column rename plan and the feature spec ran out of room and met no check. F06 F07 A01 A03 A04
  • At low effort several answers ran over their length limits, including the pull request review, the booking plan, the microservices answer, the text to the plumber and the book summary. F04 A01 A02 W01 H02
  • At low effort the cycling paragraph took the printed total on trust and gave a percentage the sources did not, and the column rename plan dropped the old column before old app versions had stopped using it. R02 A03

What max effort changed

  • Overall it scored lower at max effort: 6 prompts gained at max effort, 10 lost ground and 8 stayed the same, for about 9.4 times the cost and 11.3 times the time.
  • Every prompt that lost ground had run out of room. Thinking took 546,690 tokens across the max effort answers, against 9,654 at low effort; web pages fell from 98 to 0 and planning from 85 to 25. B01 F06 F07 A01 A03 A04
  • Where it finished, max effort helped on everyday prompts, which rose from 89 to 100: the plumber text went from 5 of 6 checks to 6 of 6, the cycling paragraph from 4 of 6 to 6 of 6. W01 R02

In short

At low effort Sonnet 5.5 scored level with Sonnet 5 at low effort, with a median answer time of 8.8 seconds and a median cost of $0.017. Overall it scored lower at max effort, because 10 of its max effort answers ran out of room.

Worth knowing

  • Each prompt was answered once at each effort, so a single answer can move a prompt's score, and small gaps between scores are treated as level.
  • Every model gets the same output allowance. An answer that runs out of room is scored on whatever it did write, so an unfinished page or program usually fails its tests.
  • Costs are what the same tokens would cost at the published API prices, not what anyone paid.

Score by kind of work

Building web pages

  • Low effort
    98
  • Max effort
    0

Code

  • Low effort
    98
  • Max effort
    71

Planning and architecture

  • Low effort
    85
  • Max effort
    25

Everyday

  • Low effort
    89
  • Max effort
    100

Every answer

Each row opens to the low and max answers side by side. Each prompt is described on the prompts page.

Building web pages

B01A landing page for a small invoicing appLow 9/9Max 0/9Max −9

What it asks for. Asks for a one-page landing site for a small invoicing app for freelancers, as a single HTML file with no libraries and nothing loaded from outside. The page needs a headline, three named features, three price plans with a monthly and yearly switch (yearly gives two months free and is quoted as one total for the year), a four-question FAQ that opens and closes, and a sign-up button. It has to read well on a phone.

Max effort passed 9 fewer checks than low.

AI answer from Sonnet 5.5 · low effort · 2 Oct 2026

9 of 9

Screenshot of the page built in this answer, at a desktop width of 1280 pixels
Our screenshot, 1280 by 800. Open it for the full length.
Screenshot of the page built in this answer, at a phone width of 375 pixels
Phone, 375 by 812.
More screenshots (2).
Screenshot of the page built in this answer, at a desktop width of 1280 pixelsScreenshot of the page built in this answer, at a phone width of 375 pixels
Show the code and the answer.

Yearly prices are 10 × monthly (two months free): Starter $90, Pro $190, Team $490.

<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Brindlebill – Invoicing for freelancers</title>
<meta name="description" content="Brindlebill is simple invoicing for freelancers. Send invoices, chase payments gently, and get paid by card or bank.">
<style>
  :root {
    --ink: #1d2433;
    --muted: #4b5563;
    --bg: #fbf9f4;
    --card: #ffffff;
    --line: #e2ddd0;
    --accent: #1f6f5c;
    --accent-dark: #175647;
    --accent-soft: #e3f1ec;
  }
  * { box-sizing: border-box; }
  html { -webkit-text-size-adjust: 100%; }
  body {
    margin: 0;
    font-family: system-ui, -apple-system, "Segoe UI", Roboto, Helvetica, Arial, sans-serif;
    font-size: 1rem;
    line-height: 1.6;
    color: var(--ink);
    background: var(--bg);
  }
  .wrap { max-width: 64rem; margin: 0 auto; padding: 0 1.25rem; }
  section { padding: 3rem 0; }
  h1, h2, h3 { line-height: 1.2; margin: 0 0 .75rem; }
  h1 { font-size: clamp(2rem, 7vw, 3.25rem); }
  h2 { font-size: clamp(1.5rem, 5vw, 2rem); text-align: center; margin-bottom: 1.5rem; }
  h3 { font-size: 1.25rem; }
  p { margin: 0 0 1rem; }
  .brand { font-weight: 800; font-size: 1.25rem; color: var(--accent); padding: 1.25rem 0 0; }
  .hero { text-align: center; padding: 3rem 0 3.5rem; }
  .pitch { font-size: 1.25rem; color: var(--muted); max-width: 36rem; margin: 0 auto 1.75rem; }
  .btn {
    display: inline-block;
    background: var(--accent);
    color: #fff;
    font: inherit;
    font-weight: 700;
    text-decoration: none;
    padding: .85rem 1.75rem;
    min-height: 44px;
    border: 0;
    border-radius: .6rem;
    cursor: pointer;
  }
  .btn:hover { background: var(--accent-dark); }
  a:focus-visible, button:focus-visible, summary:focus-visible { outline: 3px solid #f0a500; outline-offset: 2px; }
  .grid { display: grid; gap: 1rem; }
  .features { grid-template-columns: repeat(auto-fit, minmax(15rem, 1fr)); }
  .card {
    background: var(--card);
    border: 1px solid var(--line);
    border-radius: .8rem;
    padding: 1.5rem;
  }
  .card p:last-child { margin-bottom: 0; }

  .toggle {
    display: flex;
    justify-content: center;
    gap: .25rem;
    background: var(--accent-soft);
    border-radius: 999px;
    padding: .25rem;
    width: fit-content;
    margin: 0 auto 1.5rem;
  }
  .toggle button {
    font: inherit;
    font-weight: 600;
    color: var(--ink);
    background: transparent;
    border: 0;
    border-radius: 999px;
    padding: .5rem 1.25rem;
    min-height: 44px;
    cursor: pointer;
  }
  .toggle button[aria-pressed="true"] { background: var(--accent); color: #fff; }
  .plans { grid-template-columns: repeat(auto-fit, minmax(16rem, 1fr)); }
  .plan h3 { margin-bottom: .25rem; }
  .price { font-size: 2.25rem; font-weight: 800; margin: .5rem 0 0; line-height: 1.1; }
  .period { color: var(--muted); margin-bottom: 1rem; }
  .plan ul { padding-left: 1.25rem; margin: 0; }
  .plan li { margin-bottom: .4rem; }
  .plan.featured { border: 2px solid var(--accent); }
  .note { text-align: center; color: var(--muted); min-height: 1.6em; }

  .faq { max-width: 44rem; margin: 0 auto; }
  details {
    background: var(--card);
    border: 1px solid var(--line);
    border-radius: .6rem;
    margin-bottom: .75rem;
  }
  summary {
    cursor: pointer;
    font-weight: 600;
    padding: 1rem 1.25rem;
    min-height: 44px;
  }
  details p { padding: 0 1.25rem; margin: 0 0 1rem; }

  .cta { text-align: center; }
  footer { text-align: center; color: var(--muted); padding: 2rem 0; border-top: 1px solid var(--line); }
</style>
</head>
<body>
<header class="wrap"><div class="brand">Brindlebill</div></header>

<main>
  <section class="hero wrap">
    <h1>Invoicing that gets you paid.</h1>
    <p class="pitch">Brindlebill helps freelancers send professional invoices and get paid on time, without the admin.</p>
    <a class="btn" href="#signup">Start free trial</a>
  </section>

  <section class="wrap" aria-labelledby="features-h">
    <h2 id="features-h">Why freelancers choose Brindlebill</h2>
    <div class="grid features">
      <div class="card">
        <h3>Invoices in a minute</h3>
        <p>Pick a client, add your line items and send a polished invoice before your coffee cools.</p>
      </div>
      <div class="card">
        <h3>Gentle reminders</h3>
        <p>Friendly nudges go out when an invoice is overdue, so you never have to chase awkwardly.</p>
      </div>
      <div class="card">
        <h3>Card and bank payments</h3>
        <p>Clients pay straight from the invoice by card or bank transfer, and you see it the moment it lands.</p>
      </div>
    </div>
  </section>

  <section class="wrap" aria-labelledby="pricing-h">
    <h2 id="pricing-h">Pricing</h2>
    <div class="toggle" role="group" aria-label="Billing period">
      <button type="button" id="btn-monthly" aria-pressed="true">Monthly</button>
      <button type="button" id="btn-yearly" aria-pressed="false">Yearly</button>
    </div>
    <p class="note" id="note" aria-live="polite">Billed monthly.</p>
    <div class="grid plans">
      <div class="card plan">
        <h3>Starter</h3>
        <p class="price"><span data-monthly="9" data-yearly="90">$9</span></p>
        <p class="period" data-period>per month</p>
        <ul>
          <li>5 clients</li>
          <li>PDF invoices</li>
        </ul>
      </div>
      <div class="card plan featured">
        <h3>Pro</h3>
        <p class="price"><span data-monthly="19" data-yearly="190">$19</span></p>
        <p class="period" data-period>per month</p>
        <ul>
          <li>Unlimited clients</li>
          <li>Automatic reminders</li>
          <li>Your own logo</li>
        </ul>
      </div>
      <div class="card plan">
        <h3>Team</h3>
        <p class="price"><span data-monthly="49" data-yearly="490">$49</span></p>
        <p class="period" data-period>per month</p>
        <ul>
          <li>Everything in Pro</li>
          <li>5 seats</li>
          <li>A shared client list</li>
        </ul>
      </div>
    </div>
  </section>

  <section class="wrap" aria-labelledby="faq-h">
    <h2 id="faq-h">Questions, answered</h2>
    <div class="faq">
      <details>
        <summary>Can I cancel any time?</summary>
        <p>Yes. Cancel from your account page and you won't be charged again. Nothing is locked in.</p>
      </details>
      <details>
        <summary>Do my clients need an account?</summary>
        <p>No. They open a link, see the invoice and pay. They never have to sign up.</p>
      </details>
      <details>
        <summary>Which currencies can I invoice in?</summary>
        <p>You can invoice in 30 currencies, and each invoice shows the amount in your client's currency.</p>
      </details>
      <details>
        <summary>What happens when the free trial ends?</summary>
        <p>You pick a plan. If you don't, your account switches to view-only and your invoices stay safe for 90 days.</p>
      </details>
    </div>
  </section>

  <section class="wrap cta" id="signup" aria-labelledby="cta-h">
    <h2 id="cta-h">Ready when you are</h2>
    <a class="btn" href="#signup">Start free trial</a>
  </section>
</main>

<footer><div class="wrap">&copy; Brindlebill</div></footer>

<script>
  (function () {
    var monthlyBtn = document.getElementById('btn-monthly');
    var yearlyBtn = document.getElementById('btn-yearly');
    var note = document.getElementById('note');
    var prices = document.querySelectorAll('[data-monthly]');
    var periods = document.querySelectorAll('[data-period]');

    function setPeriod(yearly) {
      monthlyBtn.setAttribute('aria-pressed', String(!yearly));
      yearlyBtn.setAttribute('aria-pressed', String(yearly));
      prices.forEach(function (el) {
        el.textContent = '$' + (yearly ? el.dataset.yearly : el.dataset.monthly);
      });
      periods.forEach(function (el) {
        el.textContent = yearly ? 'per year' : 'per month';
      });
      note.textContent = yearly ? 'Billed yearly. You get two months free.' : 'Billed monthly.';
    }

    monthlyBtn.addEventListener('click', function () { setPeriod(false); });
    yearlyBtn.addEventListener('click', function () { setPeriod(true); });
  })();
</script>
</body>
</html>

Notes:

  • Yearly prices are $90, $190 and $490, shown per year, with "per year" under each.
  • The FAQ uses native <details>, so it opens and closes without JavaScript and works with keyboards and screen readers.
  • Both "Start free trial" buttons link to #signup. Point the href at your real sign-up URL when you have one.
  • The page uses a single h1, with h2 for sections and h3 for features and plans. Body text is 16px or larger.
  • The feature sentences are my own wording, since you only gave titles. Edit them as you like.
  • Floor2/2, 100% passed
  • Middle4/4, 100% passed
  • Top3/3, 100% passed
9 of 9 checks passedfloor 2/2, middle 4/4, top 3/3
  • Passed.
    The page loads with no errors in the console
    Floor Page test
  • Passed.
    Every section is there, with the exact plan names, prices and text
    Floor Page test
  • Passed.
    No sideways scroll on a 375 px phone
    Middle Page test
  • Passed.
    Nothing is loaded from outside the file
    Middle Page test
  • Passed.
    One h1, no skipped heading levels, text 16 px or larger, enough contrast
    Middle Page test
  • Passed.
    The page looks finished: clear hierarchy, pricing easy to compare, nothing broken
    Middle Read
  • Passed.
    Yearly shows $90, $190 and $490, and Monthly brings back $9, $19, $49
    Top Page test
  • Passed.
    Each FAQ question opens and closes from the keyboard
    Top Page test
  • Passed.
    Still no sideways scroll on a phone with Yearly on and every FAQ open
    Top Page test
  • Took 20 s.
  • First words after 1.6 s.
  • Wrote 3,835 tokens.
  • 4.2¢ at API prices.
All numbers.
Time to finish20 s
Time to first words1.6 s
Input tokens2
Cached tokens1,275
Output tokens3,835
Thinking tokens0
Words in the answer982
Price at API rates4.2¢

AI answer from Sonnet 5.5 · max effort · 2 Oct 2026

0 of 9

The answer was cut off at the token limit.

  • Floor0/2, 0% passed
  • Middle0/4, 0% passed
  • Top0/3, 0% passed
0 of 9 checks passedfloor 0/2, middle 0/4, top 0/3
  • Missed.
    The page loads with no errors in the consoleNo answer
    Floor Page test
  • Missed.
    Every section is there, with the exact plan names, prices and textNo answer
    Floor Page test
  • Missed.
    No sideways scroll on a 375 px phoneNo answer
    Middle Page test
  • Missed.
    Nothing is loaded from outside the fileNo answer
    Middle Page test
  • Missed.
    One h1, no skipped heading levels, text 16 px or larger, enough contrastNo answer
    Middle Page test
  • Missed.
    The page looks finished: clear hierarchy, pricing easy to compare, nothing brokenNo answer
    Middle Read
  • Missed.
    Yearly shows $90, $190 and $490, and Monthly brings back $9, $19, $49No answer
    Top Page test
  • Missed.
    Each FAQ question opens and closes from the keyboardNo answer
    Top Page test
  • Missed.
    Still no sideways scroll on a phone with Yearly on and every FAQ openNo answer
    Top Page test
  • Took 4 min 7 s.
  • Wrote 32,000 tokens, 32,000 of them thinking.
  • 32¢ at API prices.
All numbers.
Time to finish4 min 7 s
Time to first wordsnot recorded
Input tokens2
Cached tokens1,275
Output tokens32,000
Thinking tokens32,000
Words in the answer0
Price at API rates32¢
B02Build a dashboard from a picture of its designLow 7/7Max 0/7Max −7

What it asks for. Sends a picture of a designer's mockup, an admin dashboard for a small bike repair shop, and asks for it to be built as one HTML file that matches the layout, colours, spacing and text as closely as it can. The bar chart has to be drawn to scale. On a phone the page must stack into one column, hide the sidebar behind a menu button and never scroll sideways.

Max effort passed 7 fewer checks than low.

AI answer from Sonnet 5.5 · low effort · 2 Oct 2026

7 of 7

Screenshot of the page built in this answer, at a desktop width of 1280 pixels
Our screenshot, 1280 by 800. Open it for the full length.
Screenshot of the page built in this answer, at a phone width of 375 pixels
Phone, 375 by 812.
More screenshots (1).
Screenshot of the page built in this answer, at a phone width of 375 pixels
Show the code and the answer.

I've built the mockup as a single HTML file with no libraries, outside fonts or images. The chart is plain HTML/CSS. Each bar's height is its value × 4px, so the bars are drawn to scale. On screens 820px wide or narrower, everything stacks into one column and the sidebar slides in from a menu button.

<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Tarnwick Bikeworks – Overview</title>
<style>
  *{box-sizing:border-box;margin:0;padding:0}
  html,body{overflow-x:hidden}
  body{
    font-family:-apple-system,BlinkMacSystemFont,"Segoe UI",Inter,Roboto,Helvetica,Arial,sans-serif;
    background:#f1f3f8;color:#1a2233;font-size:14px;line-height:1.3;
  }
  button{font:inherit;cursor:pointer}

  /* Sidebar */
  .sidebar{
    position:fixed;top:0;left:0;bottom:0;width:240px;background:#1b2845;z-index:30;
    transition:transform .25s ease;
  }
  .brand{height:65px;display:flex;align-items:center;padding:0 24px;color:#fff;font-weight:700;font-size:16px;border-bottom:1px solid #2a3858}
  .nav{padding:23px 12px;display:flex;flex-direction:column;gap:4px}
  .nav a{
    display:flex;align-items:center;gap:12px;height:40px;padding:0 14px;border-radius:8px;
    color:#c3cde2;text-decoration:none;font-size:15px;font-weight:500;
  }
  .nav a::before{content:"";width:16px;height:16px;border:1.5px solid currentColor;border-radius:5px;flex:none;opacity:.85}
  .nav a.active{background:#2f6df0;color:#fff}
  .scrim{display:none;position:fixed;inset:0;background:rgba(10,18,35,.5);z-index:20}

  /* Main */
  .main{margin-left:240px;min-height:100vh}
  header{
    height:65px;background:#fff;border-bottom:1px solid #e3e7ef;display:flex;align-items:center;
    padding:0 32px;gap:16px;
  }
  .menu-btn{display:none;width:40px;height:40px;border:1px solid #dfe3ec;background:#fff;border-radius:8px;flex:none;align-items:center;justify-content:center}
  .menu-btn span,.menu-btn span::before,.menu-btn span::after{display:block;width:18px;height:2px;background:#1a2233;border-radius:2px;position:relative;content:""}
  .menu-btn span::before{position:absolute;top:-6px}
  .menu-btn span::after{position:absolute;top:6px}
  h1{font-size:22px;font-weight:700;letter-spacing:-.2px;flex:none}
  .spacer{flex:1}
  .search{
    width:292px;height:34px;background:#f1f3f8;border:1px solid #dfe3ec;border-radius:8px;
    display:flex;align-items:center;gap:8px;padding:0 12px;color:#5b6579;font-size:14px;min-width:0
  }
  .search svg{flex:none}
  .search input{border:0;background:transparent;font:inherit;color:#1a2233;outline:none;width:100%;min-width:0}
  .search input::placeholder{color:#5b6579}
  .avatar{width:36px;height:36px;border-radius:50%;background:#1b2845;color:#fff;font-size:12px;font-weight:700;display:flex;align-items:center;justify-content:center;flex:none;margin-left:16px}

  .content{padding:31px 32px 50px;display:flex;flex-direction:column;gap:32px}
  .card{background:#fff;border:1px solid #e6e9f1;border-radius:14px}

  .stats{display:grid;grid-template-columns:repeat(3,1fr);gap:32px}
  .stat{padding:16px 19px 0;height:109px}
  .stat .label{font-size:13px;font-weight:600;color:#5b6579}
  .stat .value{font-size:32px;font-weight:700;line-height:1.1;margin-top:4px;letter-spacing:-.3px}
  .delta{font-size:13px;font-weight:600;margin-top:6px;display:flex;align-items:center;gap:6px}
  .delta::before{content:"";border:5px solid transparent}
  .delta.up{color:#1a7a3c}
  .delta.up::before{border-bottom:8px solid #1a7a3c;border-top:0}
  .delta.down{color:#c8321f}
  .delta.down::before{border-top:8px solid #c8321f;border-bottom:0}

  .chart-card{padding:16px 23px 0;height:230px}
  .card-head{display:flex;justify-content:space-between;align-items:baseline}
  .card-head h2{font-size:15px;font-weight:700}
  .card-head .range{font-size:13px;color:#5b6579}
  .chart{position:relative;margin-top:14px;height:170px}
  .plot{position:absolute;left:0;right:0;top:0;height:146px;border-bottom:1px solid #dfe3ec}
  .plot::before{content:"";position:absolute;left:0;right:0;bottom:62px;border-top:1px solid #eceff5}
  .cols{position:absolute;inset:0;display:grid;grid-template-columns:repeat(7,1fr)}
  .col{display:flex;flex-direction:column;align-items:center;justify-content:flex-end;position:relative}
  .col .val{font-size:13px;font-weight:700;margin-bottom:5px;line-height:1}
  .bar{width:min(56px,60%);height:calc(var(--v) * 4px);background:#2f6df0;border-radius:4px 4px 0 0}
  .days{position:absolute;left:0;right:0;top:152px;display:grid;grid-template-columns:repeat(7,1fr);text-align:center;font-size:13px;color:#5b6579}

  .jobs{padding:0 0 15px}
  .jobs .card-head{padding:16px 23px 12px}
  .row{display:grid;grid-template-columns:96px 178px 230px 120px 1fr;align-items:center;padding:0 23px;height:44px;border-top:1px solid #eceff5;font-size:14px}
  .row.head{height:32px;background:#f5f7fb;border-top:0;font-size:12px;font-weight:600;color:#5b6579}
  .row .id{font-weight:700}
  .pill{display:inline-block;padding:4px 14px;border-radius:999px;font-size:12px;font-weight:700;justify-self:start;white-space:nowrap}
  .ready{background:#d8f3e1;color:#17703a}
  .prog{background:#fdebc8;color:#7a4a05}
  .wait{background:#e4defa;color:#4a2fa0}

  /* Phone / narrow */
  @media (max-width:820px){
    .sidebar{transform:translateX(-100%)}
    body.open .sidebar{transform:none}
    body.open .scrim{display:block}
    .main{margin-left:0}
    .menu-btn{display:flex}
    header{padding:0 16px;gap:12px}
    h1{font-size:19px}
    .search{width:auto;flex:1;max-width:none}
    .avatar{margin-left:0}
    .spacer{display:none}
    .content{padding:20px 16px 32px;gap:16px}
    .stats{grid-template-columns:1fr;gap:16px}
    .stat{height:auto;padding:16px 19px 16px}
    .chart-card{padding-left:12px;padding-right:12px}
    .row.head{display:none}
    .row{grid-template-columns:auto 1fr;height:auto;padding:12px 16px;row-gap:4px;column-gap:12px}
    .row>*::before{content:attr(data-label);color:#5b6579;font-size:12px;font-weight:600;margin-right:0}
    .row>*{display:contents}
    .row .cell{display:flex;gap:12px;grid-column:1 / -1;align-items:center}
    .row .cell::before{content:attr(data-label);width:60px;flex:none;color:#5b6579;font-size:12px;font-weight:600}
  }
  @media (max-width:420px){
    .search{display:none}
    header .spacer{display:block}
  }
</style>
</head>
<body>

<aside class="sidebar" id="sidebar">
  <div class="brand">Tarnwick Bikeworks</div>
  <nav class="nav">
    <a href="#" class="active">Overview</a>
    <a href="#">Jobs</a>
    <a href="#">Bookings</a>
    <a href="#">Stock</a>
    <a href="#">Reports</a>
  </nav>
</aside>
<div class="scrim" id="scrim"></div>

<div class="main">
  <header>
    <button class="menu-btn" id="menuBtn" aria-label="Open menu" aria-controls="sidebar" aria-expanded="false"><span></span></button>
    <h1>Overview</h1>
    <div class="spacer"></div>
    <label class="search">
      <svg width="16" height="16" viewBox="0 0 16 16" fill="none" stroke="#5b6579" stroke-width="1.5" stroke-linecap="round"><circle cx="7" cy="7" r="5"/><path d="M11 11l3.5 3.5"/></svg>
      <input type="text" placeholder="Search jobs and bikes">
    </label>
    <div class="avatar">TB</div>
  </header>

  <div class="content">
    <section class="stats">
      <div class="card stat">
        <div class="label">Revenue</div>
        <div class="value">$8,460</div>
        <div class="delta up">12.4% vs last week</div>
      </div>
      <div class="card stat">
        <div class="label">Jobs completed</div>
        <div class="value">130</div>
        <div class="delta up">8.3% vs last week</div>
      </div>
      <div class="card stat">
        <div class="label">Open bookings</div>
        <div class="value">24</div>
        <div class="delta down">6.0% vs last week</div>
      </div>
    </section>

    <section class="card chart-card">
      <div class="card-head"><h2>Completed per day</h2><span class="range">Last 7 days</span></div>
      <div class="chart">
        <div class="plot"><div class="cols">
          <div class="col"><span class="val">14</span><div class="bar" style="--v:14"></div></div>
          <div class="col"><span class="val">22</span><div class="bar" style="--v:22"></div></div>
          <div class="col"><span class="val">18</span><div class="bar" style="--v:18"></div></div>
          <div class="col"><span class="val">27</span><div class="bar" style="--v:27"></div></div>
          <div class="col"><span class="val">31</span><div class="bar" style="--v:31"></div></div>
          <div class="col"><span class="val">12</span><div class="bar" style="--v:12"></div></div>
          <div class="col"><span class="val">6</span><div class="bar" style="--v:6"></div></div>
        </div></div>
        <div class="days"><span>Mon</span><span>Tue</span><span>Wed</span><span>Thu</span><span>Fri</span><span>Sat</span><span>Sun</span></div>
      </div>
    </section>

    <section class="card jobs">
      <div class="card-head"><h2>Recent jobs</h2></div>
      <div class="row head"><span>Job</span><span>Bike</span><span>Work</span><span>Due</span><span>Status</span></div>

      <div class="row"><span class="cell id" data-label="Job">#2041</span><span class="cell" data-label="Bike">Trail hardtail</span><span class="cell" data-label="Work">Gear tune and chain</span><span class="cell" data-label="Due">12 Mar</span><span class="cell" data-label="Status"><span class="pill ready">Ready for pickup</span></span></div>
      <div class="row"><span class="cell id" data-label="Job">#2040</span><span class="cell" data-label="Bike">Commuter 700c</span><span class="cell" data-label="Work">Brake pads, both wheels</span><span class="cell" data-label="Due">12 Mar</span><span class="cell" data-label="Status"><span class="pill prog">In progress</span></span></div>
      <div class="row"><span class="cell id" data-label="Job">#2039</span><span class="cell" data-label="Bike">Folding bike</span><span class="cell" data-label="Work">New rear tyre</span><span class="cell" data-label="Due">13 Mar</span><span class="cell" data-label="Status"><span class="pill wait">Waiting for parts</span></span></div>
      <div class="row"><span class="cell id" data-label="Job">#2038</span><span class="cell" data-label="Bike">Road racer</span><span class="cell" data-label="Work">Full service</span><span class="cell" data-label="Due">14 Mar</span><span class="cell" data-label="Status"><span class="pill prog">In progress</span></span></div>
      <div class="row"><span class="cell id" data-label="Job">#2037</span><span class="cell" data-label="Bike">Cargo trike</span><span class="cell" data-label="Work">Wheel true, front</span><span class="cell" data-label="Due">14 Mar</span><span class="cell" data-label="Status"><span class="pill ready">Ready for pickup</span></span></div>
    </section>
  </div>
</div>

<script>
  const body=document.body, btn=document.getElementById('menuBtn');
  function setMenu(open){
    body.classList.toggle('open',open);
    btn.setAttribute('aria-expanded',open);
    btn.setAttribute('aria-label',open?'Close menu':'Open menu');
  }
  btn.addEventListener('click',()=>setMenu(!body.classList.contains('open')));
  document.getElementById('scrim').addEventListener('click',()=>setMenu(false));
  document.addEventListener('keydown',e=>{if(e.key==='Escape')setMenu(false)});
  document.querySelectorAll('.nav a').forEach(a=>a.addEventListener('click',e=>{e.preventDefault();setMenu(false)}));
</script>
</body>
</html>

Notes:

  • Chart: each bar's height is value × 4px, so 31 gives 124px and 6 gives 24px. The faint gridline sits at the halfway mark.
  • Colours and spacing: these are estimated by eye from the image, since I can't sample exact values. If your designer has a spec, the colours are easy to swap.
  • Phone layout: the table rows become small stacked cards with a label beside each value. On very narrow screens (420px or less) the search box is hidden to leave room for the menu button and title.
  • Menu: the sidebar closes when you tap the dimmed background, press Escape, or tap a nav link.
  • Nav icons: the mockup shows plain rounded squares, so I drew those rather than guessing at real icons.
  • Floor2/2, 100% passed
  • Middle2/2, 100% passed
  • Top3/3, 100% passed
7 of 7 checks passedfloor 2/2, middle 2/2, top 3/3
  • Passed.
    Page loads with no console errors and nothing fetched from outside
    Floor Page test
  • Passed.
    Every word and number in the picture is on the page
    Floor Page test
  • Passed.
    Wide screen: full-height left sidebar, three cards in a row, table cells in the right columns
    Middle Page test
  • Passed.
    Sidebar, active item, accent blue and the three status pills match the picture
    Middle Page test
  • Passed.
    Chart bars run Mon to Sun with heights in proportion to their values
    Top Page test
  • Passed.
    On a phone: no sideways scroll, cards stacked, sidebar hidden until a menu button opens it
    Top Page test
  • Passed.
    Side by side with the picture, a designer would accept it as a faithful build
    Top Read
  • Took 42 s.
  • First words after 15 s.
  • Wrote 6,618 tokens, 1,309 of them thinking.
  • 7.2¢ at API prices.
All numbers.
Time to finish42 s
Time to first words15 s
Input tokens2
Cached tokens2,330
Output tokens6,618
Thinking tokens1,309
Words in the answer709
Price at API rates7.2¢

AI answer from Sonnet 5.5 · max effort · 2 Oct 2026

0 of 7

The answer was cut off at the token limit.

  • Floor0/2, 0% passed
  • Middle0/2, 0% passed
  • Top0/3, 0% passed
0 of 7 checks passedfloor 0/2, middle 0/2, top 0/3
  • Missed.
    Page loads with no console errors and nothing fetched from outsideNo answer
    Floor Page test
  • Missed.
    Every word and number in the picture is on the pageNo answer
    Floor Page test
  • Missed.
    Wide screen: full-height left sidebar, three cards in a row, table cells in the right columnsNo answer
    Middle Page test
  • Missed.
    Sidebar, active item, accent blue and the three status pills match the pictureNo answer
    Middle Page test
  • Missed.
    Chart bars run Mon to Sun with heights in proportion to their valuesNo answer
    Top Page test
  • Missed.
    On a phone: no sideways scroll, cards stacked, sidebar hidden until a menu button opens itNo answer
    Top Page test
  • Missed.
    Side by side with the picture, a designer would accept it as a faithful buildNo answer
    Top Read
  • Took 4 min 7 s.
  • Wrote 32,000 tokens, 32,000 of them thinking.
  • 33¢ at API prices.
All numbers.
Time to finish4 min 7 s
Time to first wordsnot recorded
Input tokens2
Cached tokens2,330
Output tokens32,000
Thinking tokens32,000
Words in the answer0
Price at API rates33¢
B03A kanban board that survives a reloadLow 8/9Max 0/9Max −8

What it asks for. Asks for a personal kanban board in one HTML file with three columns: to do, doing and done. Cards are added by typing a title, can be dragged between and within columns, edited and deleted, and each column shows how many cards it holds. The board must survive a reload, work from the keyboard with the arrow keys, show titles that contain code exactly as typed, and fit a phone screen.

Max effort passed 8 fewer checks than low.

AI answer from Sonnet 5.5 · low effort · 2 Oct 2026

8 of 9

Screenshot of the page built in this answer, at a desktop width of 1280 pixels
Our screenshot, 1280 by 800. Open it for the full length.
Show the code and the answer.
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Kanban</title>
<style>
  * { box-sizing: border-box; }
  html, body { margin: 0; overflow-x: hidden; }
  body { font-family: system-ui, sans-serif; background: #f0f2f5; color: #1c1e21; padding: 12px; }
  h1 { font-size: 1.3rem; margin: 0 0 10px; }
  form.add { display: flex; gap: 8px; margin-bottom: 12px; }
  form.add input { flex: 1; min-width: 0; }
  input[type=text] { font: inherit; padding: 8px; border: 1px solid #999; border-radius: 6px; }
  button { font: inherit; padding: 6px 10px; border: 1px solid #888; border-radius: 6px; background: #fff; cursor: pointer; }
  button:hover { background: #e8eaee; }
  .board { display: grid; grid-template-columns: repeat(3, minmax(0, 1fr)); gap: 12px; align-items: start; }
  .column { background: #dfe3e8; border-radius: 8px; padding: 8px; min-width: 0; }
  .column h2 { font-size: 1rem; margin: 0 0 8px; }
  .list { min-height: 48px; display: flex; flex-direction: column; gap: 8px; }
  .card { background: #fff; border: 1px solid #bbb; border-radius: 6px; padding: 8px; min-width: 0; user-select: none; -webkit-user-select: none; }
  .card:focus-visible, button:focus-visible, input:focus-visible { outline: 3px solid #1a73e8; outline-offset: 1px; }
  .top { display: flex; gap: 8px; align-items: flex-start; }
  .handle { touch-action: none; cursor: grab; padding: 0 4px; color: #666; font-size: 1.2rem; line-height: 1.2; }
  .title { flex: 1; min-width: 0; overflow-wrap: anywhere; white-space: pre-wrap; }
  .actions { display: flex; gap: 6px; margin-top: 8px; flex-wrap: wrap; }
  .edit { display: flex; gap: 6px; flex-wrap: wrap; flex: 1; min-width: 0; }
  .edit input { flex: 1 1 100%; min-width: 0; }
  .placeholder { border: 2px dashed #1a73e8; border-radius: 6px; background: rgba(26,115,232,.08); }
  .ghost { position: fixed; z-index: 10; pointer-events: none; opacity: .85; box-shadow: 0 4px 14px rgba(0,0,0,.3); }
  .sr { position: absolute; left: -9999px; }
  @media (max-width: 640px) { .board { grid-template-columns: minmax(0, 1fr); } }
</style>
</head>
<body>
<h1>My tasks</h1>
<form class="add" id="addForm">
  <input type="text" id="newTitle" placeholder="New card title" aria-label="New card title" autocomplete="off">
  <button type="submit">Add card</button>
</form>
<div class="board" id="board"></div>
<div class="sr" id="status" role="status" aria-live="polite"></div>

<script>
(function () {
  var KEY = "kanban-board-v1";
  var COLS = [["todo", "To do"], ["doing", "Doing"], ["done", "Done"]];
  var state = load();
  var editingId = null;
  var board = document.getElementById("board");
  var statusEl = document.getElementById("status");

  function load() {
    var s = { cols: { todo: [], doing: [], done: [] }, next: 1 };
    try {
      var raw = JSON.parse(localStorage.getItem(KEY));
      if (raw && raw.cols) {
        COLS.forEach(function (c) {
          var arr = Array.isArray(raw.cols[c[0]]) ? raw.cols[c[0]] : [];
          s.cols[c[0]] = arr.filter(function (x) {
            return x && typeof x.title === "string" && typeof x.id === "number";
          });
        });
        var max = 0;
        COLS.forEach(function (c) { s.cols[c[0]].forEach(function (x) { if (x.id > max) max = x.id; }); });
        s.next = Math.max(max + 1, typeof raw.next === "number" ? raw.next : 1);
      }
    } catch (e) {}
    return s;
  }
  function save() {
    try { localStorage.setItem(KEY, JSON.stringify(state)); } catch (e) {}
  }
  function find(id) {
    for (var i = 0; i < COLS.length; i++) {
      var arr = state.cols[COLS[i][0]];
      for (var j = 0; j < arr.length; j++) if (arr[j].id === id) return { col: i, idx: j };
    }
    return null;
  }
  function focusCard(id) {
    var el = board.querySelector('.card[data-id="' + id + '"]');
    if (el) el.focus();
  }
  function say(t) { statusEl.textContent = t; }

  function el(tag, cls, text) {
    var e = document.createElement(tag);
    if (cls) e.className = cls;
    if (text !== undefined) e.textContent = text;
    return e;
  }

  function render() {
    board.textContent = "";
    COLS.forEach(function (c) {
      var cards = state.cols[c[0]];
      var col = el("section", "column");
      col.dataset.col = c[0];
      col.appendChild(el("h2", "", c[1] + " " + cards.length));
      var list = el("div", "list");
      cards.forEach(function (card) { list.appendChild(renderCard(card)); });
      col.appendChild(list);
      board.appendChild(col);
    });
    if (editingId !== null) {
      var inp = board.querySelector(".edit input");
      if (inp) { inp.focus(); inp.setSelectionRange(inp.value.length, inp.value.length); }
    }
  }

  function renderCard(card) {
    var d = el("div", "card");
    d.dataset.id = card.id;
    if (editingId === card.id) {
      var form = el("form", "edit");
      var inp = el("input");
      inp.type = "text";
      inp.value = card.title;
      inp.setAttribute("aria-label", "Edit title");
      var saveB = el("button", "", "Save"); saveB.type = "submit";
      var cancelB = el("button", "", "Cancel"); cancelB.type = "button";
      form.appendChild(inp); form.appendChild(saveB); form.appendChild(cancelB);
      form.addEventListener("submit", function (e) {
        e.preventDefault();
        var v = inp.value.trim();
        if (!v) { inp.focus(); return; }
        card.title = v;
        editingId = null;
        save(); render(); focusCard(card.id);
      });
      cancelB.addEventListener("click", function () { editingId = null; render(); focusCard(card.id); });
      inp.addEventListener("keydown", function (e) {
        if (e.key === "Escape") { editingId = null; render(); focusCard(card.id); }
      });
      d.appendChild(form);
      return d;
    }
    d.tabIndex = 0;
    d.setAttribute("aria-label", card.title + ". Use arrow keys to move.");
    var top = el("div", "top");
    top.appendChild(el("span", "handle", "\u2807"));
    top.appendChild(el("div", "title", card.title));
    d.appendChild(top);
    var act = el("div", "actions");
    var eb = el("button", "", "Edit"); eb.type = "button";
    var db = el("button", "", "Delete"); db.type = "button";
    eb.addEventListener("click", function () { editingId = card.id; render(); });
    db.addEventListener("click", function () {
      var p = find(card.id);
      if (p) { state.cols[COLS[p.col][0]].splice(p.idx, 1); }
      if (editingId === card.id) editingId = null;
      save(); render();
    });
    act.appendChild(eb); act.appendChild(db);
    d.appendChild(act);

    d.addEventListener("keydown", function (e) {
      if (e.target !== d) return;
      var dx = 0, dy = 0;
      if (e.key === "ArrowLeft") dx = -1;
      else if (e.key === "ArrowRight") dx = 1;
      else if (e.key === "ArrowUp") dy = -1;
      else if (e.key === "ArrowDown") dy = 1;
      else return;
      e.preventDefault();
      var p = find(card.id);
      if (!p) return;
      var from = state.cols[COLS[p.col][0]];
      if (dx) {
        var nc = p.col + dx;
        if (nc < 0 || nc >= COLS.length) return;
        var to = state.cols[COLS[nc][0]];
        from.splice(p.idx, 1);
        to.splice(Math.min(p.idx, to.length), 0, card);
        say("Moved to " + COLS[nc][1]);
      } else {
        var ni = p.idx + dy;
        if (ni < 0 || ni >= from.length) return;
        from.splice(p.idx, 1);
        from.splice(ni, 0, card);
        say("Moved to position " + (ni + 1));
      }
      save(); render(); focusCard(card.id);
    });

    d.addEventListener("pointerdown", function (e) { startPointer(e, d, card); });
    return d;
  }

  // Pointer-based dragging (mouse, touch, pen)
  function startPointer(e, cardEl, card) {
    if (e.button !== 0) return;
    if (e.target.closest("button, input")) return;
    if (e.pointerType === "touch" && !e.target.closest(".handle")) return;
    var sx = e.clientX, sy = e.clientY, dragging = false;
    var ghost, placeholder, rect, targetCol = null, targetIdx = 0;

    function begin() {
      dragging = true;
      rect = cardEl.getBoundingClientRect();
      ghost = cardEl.cloneNode(true);
      ghost.classList.add("ghost");
      ghost.style.width = rect.width + "px";
      document.body.appendChild(ghost);
      placeholder = el("div", "placeholder");
      placeholder.style.height = rect.height + "px";
      cardEl.parentNode.insertBefore(placeholder, cardEl);
      cardEl.style.display = "none";
    }
    function move(ev) {
      if (!dragging) {
        if (Math.abs(ev.clientX - sx) + Math.abs(ev.clientY - sy) < 6) return;
        begin();
      }
      ev.preventDefault();
      ghost.style.left = (ev.clientX - (sx - rect.left)) + "px";
      ghost.style.top = (ev.clientY - (sy - rect.top)) + "px";
      var under = document.elementFromPoint(ev.clientX, ev.clientY);
      var colEl = under && under.closest(".column");
      if (!colEl) return;
      var list = colEl.querySelector(".list");
      var others = Array.prototype.filter.call(list.querySelectorAll(".card"), function (c) { return c !== cardEl; });
      var idx = 0;
      others.forEach(function (c) {
        var r = c.getBoundingClientRect();
        if (ev.clientY > r.top + r.height / 2) idx++;
      });
      list.insertBefore(placeholder, others[idx] || null);
      targetCol = colEl.dataset.col;
      targetIdx = idx;
    }
    function cleanup() {
      document.removeEventListener("pointermove", move);
      document.removeEventListener("pointerup", up);
      document.removeEventListener("pointercancel", cancel);
      document.removeEventListener("keydown", key);
      if (dragging) { ghost.remove(); placeholder.remove(); cardEl.style.display = ""; }
    }
    function up() {
      var was = dragging;
      cleanup();
      if (was && targetCol) {
        var p = find(card.id);
        if (p) {
          state.cols[COLS[p.col][0]].splice(p.idx, 1);
          var arr = state.cols[targetCol];
          arr.splice(Math.min(targetIdx, arr.length), 0, card);
          save(); render(); focusCard(card.id);
        }
      }
    }
    function cancel() { cleanup(); }
    function key(ev) { if (ev.key === "Escape") cleanup(); }
    document.addEventListener("pointermove", move, { passive: false });
    document.addEventListener("pointerup", up);
    document.addEventListener("pointercancel", cancel);
    document.addEventListener("keydown", key);
  }

  document.getElementById("addForm").addEventListener("submit", function (e) {
    e.preventDefault();
    var input = document.getElementById("newTitle");
    var v = input.value.trim();
    if (!v) { input.value = ""; return; }
    state.cols.todo.push({ id: state.next++, title: v });
    input.value = "";
    save(); render();
    input.focus();
  });

  render();
})();
</script>
</body>
</html>

Notes on how it behaves:

  • Titles are plain text. They are only ever set with textContent or an input's value, so <div> and a & b show exactly as typed. Long titles wrap, even without spaces.
  • Phone layout. At 640px wide and below the columns stack vertically, so there is no sideways scrolling at 375px.
  • Mouse and touch dragging. Drag a card from anywhere with the mouse. On a touch screen, drag by the ⠇ handle at the left of the card so normal scrolling still works. Esc cancels a drag.
  • Keyboard. Tab onto a card, then use the arrow keys. Left and right move it to the neighbouring column, keeping its position where possible. Up and down move it within the column. The card keeps focus after each move.
  • Editing. Enter or Save confirms. Esc or Cancel discards the change. An empty title is not saved.
  • Saving. The board is saved to localStorage after every change and restored on reload. It starts empty. Because localStorage is per browser and per file location, opening the file from a different path or browser shows a separate board.
  • Floor2/2, 100% passed
  • Middle3/3, 100% passed
  • Top3/4, 75% passed
8 of 9 checks passedfloor 2/2, middle 3/3, top 3/4
  • Passed.
    The page loads with no errors and has the three named columns
    Floor Page test
  • Passed.
    Adding a card works with Enter and with the Add card button
    Floor Page test
  • Passed.
    Empty or space-only titles are ignored
    Middle Page test
  • Passed.
    The counts update, and Edit and Delete work
    Middle Page test
  • Passed.
    After a reload the board is exactly as it was
    Middle Page test
  • Passed.
    Dragging moves cards between columns and reorders them, and it survives a reload
    Top Page test
  • Passed.
    A card can be moved with the arrow keys and keeps focus
    Top Page test
  • Passed.
    Titles like <b>hi</b> & <img ...> show exactly as typed and run nothing
    Top Page test
  • Missed.
    No sideways scroll on a 375 px phone, even with very long titlestest group failed
    Top Page test
  • Took 42 s.
  • First words after 19 s.
  • Wrote 6,851 tokens, 1,850 of them thinking.
  • 7.1¢ at API prices.
All numbers.
Time to finish42 s
Time to first words19 s
Input tokens2
Cached tokens1,172
Output tokens6,851
Thinking tokens1,850
Words in the answer1,446
Price at API rates7.1¢

AI answer from Sonnet 5.5 · max effort · 2 Oct 2026

0 of 9

The answer was cut off at the token limit.

  • Floor0/2, 0% passed
  • Middle0/3, 0% passed
  • Top0/4, 0% passed
0 of 9 checks passedfloor 0/2, middle 0/3, top 0/4
  • Missed.
    The page loads with no errors and has the three named columnsNo answer
    Floor Page test
  • Missed.
    Adding a card works with Enter and with the Add card buttonNo answer
    Floor Page test
  • Missed.
    Empty or space-only titles are ignoredNo answer
    Middle Page test
  • Missed.
    The counts update, and Edit and Delete workNo answer
    Middle Page test
  • Missed.
    After a reload the board is exactly as it wasNo answer
    Middle Page test
  • Missed.
    Dragging moves cards between columns and reorders them, and it survives a reloadNo answer
    Top Page test
  • Missed.
    A card can be moved with the arrow keys and keeps focusNo answer
    Top Page test
  • Missed.
    Titles like <b>hi</b> & <img ...> show exactly as typed and run nothingNo answer
    Top Page test
  • Missed.
    No sideways scroll on a 375 px phone, even with very long titlesNo answer
    Top Page test
  • Took 4 min 9 s.
  • Wrote 32,000 tokens, 32,000 of them thinking.
  • 32¢ at API prices.
All numbers.
Time to finish4 min 9 s
Time to first wordsnot recorded
Input tokens2
Cached tokens1,172
Output tokens32,000
Thinking tokens32,000
Words in the answer0
Price at API rates32¢
B04Minesweeper in one HTML fileLow 9/9Max 0/9Max −9

What it asks for. Asks for the game Minesweeper in one HTML file: a 9 by 9 board with 10 mines, flags on right click, a mines-left counter, a first square that is never a mine, empty areas that open by themselves, win and lose messages and a new game button. Every square has to be a button so the game works from the keyboard, and the board has to fit a phone.

Max effort passed 9 fewer checks than low.

AI answer from Sonnet 5.5 · low effort · 2 Oct 2026

9 of 9

Screenshot of the page built in this answer, at a desktop width of 1280 pixels
Our screenshot, 1280 by 800. Open it for the full length.
Screenshot of the page built in this answer, at a phone width of 375 pixels
Phone, 375 by 812.
More screenshots (2).
Screenshot of the page built in this answer, at a desktop width of 1280 pixelsScreenshot of the page built in this answer, at a phone width of 375 pixels
Show the code and the answer.
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Minesweeper</title>
<style>
  * { box-sizing: border-box; }
  body {
    margin: 0;
    padding: 12px;
    font-family: system-ui, sans-serif;
    background: #eceff1;
    display: flex;
    justify-content: center;
  }
  #game {
    width: min(100%, 360px);
  }
  h1 { font-size: 1.4rem; margin: 0 0 8px; text-align: center; }
  #bar {
    display: flex;
    justify-content: space-between;
    align-items: center;
    margin-bottom: 8px;
    gap: 8px;
  }
  #counter { font-weight: bold; font-size: 1.1rem; }
  #newgame {
    min-height: 36px;
    padding: 0 14px;
    font-size: 1rem;
    cursor: pointer;
  }
  #message {
    min-height: 1.6em;
    text-align: center;
    font-weight: bold;
    font-size: 1.2rem;
    margin-bottom: 8px;
  }
  #board {
    display: grid;
    grid-template-columns: repeat(9, minmax(0, 1fr));
    gap: 2px;
    background: #90a4ae;
    padding: 2px;
    touch-action: manipulation;
  }
  .cell {
    aspect-ratio: 1;
    min-width: 32px;
    min-height: 32px;
    padding: 0;
    margin: 0;
    font-size: 1.1rem;
    font-weight: bold;
    line-height: 1;
    border: 2px outset #fafafa;
    background: #cfd8dc;
    cursor: pointer;
    display: flex;
    align-items: center;
    justify-content: center;
    user-select: none;
    -webkit-user-select: none;
  }
  .cell.open {
    border: 1px solid #90a4ae;
    background: #f5f5f5;
    cursor: default;
  }
  .cell.mine { background: #ef9a9a; }
  .cell:focus-visible { outline: 3px solid #1565c0; outline-offset: -2px; }
  .n1 { color: #1976d2; } .n2 { color: #388e3c; } .n3 { color: #d32f2f; }
  .n4 { color: #7b1fa2; } .n5 { color: #8d6e63; } .n6 { color: #00838f; }
  .n7 { color: #212121; } .n8 { color: #757575; }
  #help { font-size: .8rem; color: #455a64; margin-top: 8px; text-align: center; }
</style>
</head>
<body>
<div id="game">
  <h1>Minesweeper</h1>
  <div id="bar">
    <div id="counter">Mines left: <span id="mines-left">10</span></div>
    <button id="newgame" type="button">New game</button>
  </div>
  <div id="message" role="status" aria-live="polite"></div>
  <div id="board" role="group" aria-label="Minesweeper board"></div>
  <div id="help">Click / Enter / Space to open. Right click (or F) to flag.</div>
</div>

<script>
(function () {
  var SIZE = 9, MINES = 10, TOTAL = SIZE * SIZE;
  var boardEl = document.getElementById('board');
  var leftEl = document.getElementById('mines-left');
  var msgEl = document.getElementById('message');
  var cells = [];
  var mine, adj, opened, flagged, placed, over, openCount, flagCount;

  // Build buttons once
  for (var i = 0; i < TOTAL; i++) {
    var b = document.createElement('button');
    b.type = 'button';
    b.className = 'cell';
    b.dataset.i = i;
    boardEl.appendChild(b);
    cells.push(b);
  }

  boardEl.addEventListener('click', function (e) {
    var b = e.target.closest('.cell');
    if (b) openCell(+b.dataset.i);
  });
  boardEl.addEventListener('contextmenu', function (e) {
    e.preventDefault();
    var b = e.target.closest('.cell');
    if (b) toggleFlag(+b.dataset.i);
  });
  boardEl.addEventListener('keydown', function (e) {
    if (e.key === 'f' || e.key === 'F') {
      var b = e.target.closest('.cell');
      if (b) { e.preventDefault(); toggleFlag(+b.dataset.i); }
    }
  });
  document.getElementById('newgame').addEventListener('click', newGame);

  function neighbours(i) {
    var r = Math.floor(i / SIZE), c = i % SIZE, out = [];
    for (var dr = -1; dr <= 1; dr++) {
      for (var dc = -1; dc <= 1; dc++) {
        if (!dr && !dc) continue;
        var rr = r + dr, cc = c + dc;
        if (rr >= 0 && rr < SIZE && cc >= 0 && cc < SIZE) out.push(rr * SIZE + cc);
      }
    }
    return out;
  }

  function placeMines(safe) {
    var pool = [];
    for (var i = 0; i < TOTAL; i++) if (i !== safe) pool.push(i);
    for (var k = pool.length - 1; k > 0; k--) {
      var j = Math.floor(Math.random() * (k + 1));
      var t = pool[k]; pool[k] = pool[j]; pool[j] = t;
    }
    for (var m = 0; m < MINES; m++) mine[pool[m]] = true;
    for (var x = 0; x < TOTAL; x++) {
      var n = 0;
      neighbours(x).forEach(function (y) { if (mine[y]) n++; });
      adj[x] = n;
    }
    placed = true;
  }

  function openCell(i) {
    if (over || opened[i] || flagged[i]) return;
    if (!placed) placeMines(i);
    if (mine[i]) { lose(i); return; }
    var stack = [i];
    while (stack.length) {
      var cur = stack.pop();
      if (opened[cur] || flagged[cur] || mine[cur]) continue;
      opened[cur] = true;
      openCount++;
      render(cur);
      if (adj[cur] === 0) {
        neighbours(cur).forEach(function (n) {
          if (!opened[n] && !flagged[n]) stack.push(n);
        });
      }
    }
    if (openCount === TOTAL - MINES) {
      over = true;
      msgEl.textContent = 'You win!';
    }
  }

  function toggleFlag(i) {
    if (over || opened[i]) return;
    flagged[i] = !flagged[i];
    flagCount += flagged[i] ? 1 : -1;
    render(i);
    updateCounter();
  }

  function lose(hit) {
    over = true;
    msgEl.textContent = 'Game over';
    for (var i = 0; i < TOTAL; i++) {
      if (mine[i]) {
        cells[i].textContent = '💣';
        cells[i].className = 'cell open mine';
        cells[i].setAttribute('aria-label', 'Mine');
      }
    }
  }

  function render(i) {
    var b = cells[i];
    var r = Math.floor(i / SIZE) + 1, c = (i % SIZE) + 1;
    var pos = 'Row ' + r + ' column ' + c + ': ';
    if (opened[i]) {
      b.className = 'cell open' + (adj[i] ? ' n' + adj[i] : '');
      b.textContent = adj[i] ? adj[i] : '';
      b.setAttribute('aria-label', pos + (adj[i] ? adj[i] + ' adjacent mines' : 'empty'));
    } else if (flagged[i]) {
      b.className = 'cell';
      b.textContent = '🚩';
      b.setAttribute('aria-label', pos + 'flagged');
    } else {
      b.className = 'cell';
      b.textContent = '';
      b.setAttribute('aria-label', pos + 'unopened');
    }
  }

  function updateCounter() { leftEl.textContent = MINES - flagCount; }

  function newGame() {
    mine = []; adj = []; opened = []; flagged = [];
    for (var i = 0; i < TOTAL; i++) {
      mine.push(false); adj.push(0); opened.push(false); flagged.push(false);
    }
    placed = false; over = false; openCount = 0; flagCount = 0;
    msgEl.textContent = '';
    for (var j = 0; j < TOTAL; j++) render(j);
    updateCounter();
  }

  newGame();
})();
</script>
</body>
</html>

Notes:

  • Squares are real <button>s, so Tab/arrow focus handling and Enter/Space work natively. I also added an F key to flag the focused square, since keyboard users have no right click.
  • The counter is "Mines left" and can go below 0 if you place more than 10 flags.
  • After a win or loss the board ignores clicks and flags until you press "New game".
  • Mines are placed on your first open, so that square is never a mine.
  • Floor2/2, 100% passed
  • Middle3/3, 100% passed
  • Top4/4, 100% passed
9 of 9 checks passedfloor 2/2, middle 3/3, top 4/4
  • Passed.
    81 square buttons in a 9 by 9 grid, Mines left starts at 10
    Floor Page test
  • Passed.
    A first click never lands on a mine
    Floor Page test
  • Passed.
    Numbers are right and an empty square opens its neighbours
    Middle Page test
  • Passed.
    Right click flags and unflags, the counter follows, a flagged square stays shut
    Middle Page test
  • Passed.
    The squares work from the keyboard
    Middle Page test
  • Passed.
    Hitting a mine says Game over, shows all 10 mines as 💣 and stops the board
    Top Page test
  • Passed.
    Opening every safe square shows You win! (and not before)
    Top Page test
  • Passed.
    New game resets the board, counter and message
    Top Page test
  • Passed.
    Fits a 375 px phone with squares of at least 32 px, no sideways scroll
    Top Page test
  • Took 23 s.
  • First words after 7.8 s.
  • Wrote 3,877 tokens, 669 of them thinking.
  • 4.2¢ at API prices.
All numbers.
Time to finish23 s
Time to first words7.8 s
Input tokens2
Cached tokens1,114
Output tokens3,877
Thinking tokens669
Words in the answer952
Price at API rates4.2¢

AI answer from Sonnet 5.5 · max effort · 2 Oct 2026

0 of 9

The answer was cut off at the token limit.

  • Floor0/2, 0% passed
  • Middle0/3, 0% passed
  • Top0/4, 0% passed
0 of 9 checks passedfloor 0/2, middle 0/3, top 0/4
  • Missed.
    81 square buttons in a 9 by 9 grid, Mines left starts at 10No answer
    Floor Page test
  • Missed.
    A first click never lands on a mineNo answer
    Floor Page test
  • Missed.
    Numbers are right and an empty square opens its neighboursNo answer
    Middle Page test
  • Missed.
    Right click flags and unflags, the counter follows, a flagged square stays shutNo answer
    Middle Page test
  • Missed.
    The squares work from the keyboardNo answer
    Middle Page test
  • Missed.
    Hitting a mine says Game over, shows all 10 mines as 💣 and stops the boardNo answer
    Top Page test
  • Missed.
    Opening every safe square shows You win! (and not before)No answer
    Top Page test
  • Missed.
    New game resets the board, counter and messageNo answer
    Top Page test
  • Missed.
    Fits a 375 px phone with squares of at least 32 px, no sideways scrollNo answer
    Top Page test
  • Took 4 min 5 s.
  • Wrote 32,000 tokens, 32,000 of them thinking.
  • 32¢ at API prices.
All numbers.
Time to finish4 min 5 s
Time to first wordsnot recorded
Input tokens2
Cached tokens1,114
Output tokens32,000
Thinking tokens32,000
Words in the answer0
Price at API rates32¢
B05Fix three layout bugs on a garden pageLow 7/7Max 0/7Max −7

What it asks for. Pastes a community garden club's web page and describes three problems on phones: the page scrolls sideways, section headings end up hidden under the sticky header after tapping the menu links, and three event cards are squashed instead of sitting one per row. It asks for the whole page back with all three fixed, leaving the laptop layout as it is and the header still sticky.

Max effort passed 7 fewer checks than low.

AI answer from Sonnet 5.5 · low effort · 2 Oct 2026

7 of 7

Screenshot of the page built in this answer, at a desktop width of 1280 pixels
Our screenshot, 1280 by 800. Open it for the full length.
Screenshot of the page built in this answer, at a phone width of 375 pixels
Phone, 375 by 812.
Show the code and the answer.

Here's what was causing each problem and the fix:

  1. Sideways scrolling. Two things were wider than a phone screen. The fees table has a fixed width: 720px, and the long email address can't wrap. I made the table fill its container and scroll inside its own wrapper if it's still too wide. I also let the email address break across lines. On a laptop the table now fills the 960px content width instead of being 720px, so to keep it looking identical I set max-width: 720px on it.
  2. Headings hidden under the sticky header. I added scroll-margin-top to the sections so the browser stops short of the header. The header can be taller on a phone because the nav wraps, so I used a generous value.
  3. Squashed cards. I added a media query so the cards go to one column on narrow screens. Laptop layout is unchanged.
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Thistlebank Allotments</title>
<style>
  * { box-sizing: border-box; }
  body {
    margin: 0;
    font-family: system-ui, -apple-system, "Segoe UI", Helvetica, Arial, sans-serif;
    line-height: 1.55;
    color: #23301f;
    background: #f7f4ea;
  }
  .wrap { max-width: 960px; margin: 0 auto; padding: 0 1rem; }
  .site-header {
    position: sticky;
    top: 0;
    z-index: 10;
    background: #2f5d34;
    color: #fff;
    box-shadow: 0 1px 6px rgba(0, 0, 0, .25);
  }
  .site-header .wrap {
    display: flex;
    flex-wrap: wrap;
    align-items: center;
    justify-content: space-between;
    gap: .25rem 1.5rem;
    padding-top: .9rem;
    padding-bottom: .9rem;
  }
  .brand { font-size: 1.15rem; font-weight: 700; letter-spacing: .01em; }
  .site-header nav { display: flex; flex-wrap: wrap; gap: .25rem 1.1rem; }
  .site-header a { color: #fff; text-decoration: none; font-weight: 500; }
  .site-header a:hover { text-decoration: underline; }
  .hero { background: #dfe9c8; padding: 4rem 0 3.5rem; }
  .hero h1 { margin: 0 0 .5rem; font-size: 2.4rem; line-height: 1.15; color: #1f4424; }
  .hero p { margin: 0; max-width: 34rem; font-size: 1.15rem; }
  section.block { padding: 2rem 0; scroll-margin-top: 6rem; }
  section.block h2 { margin: 0 0 1rem; font-size: 1.7rem; color: #1f4424; }
  .cards {
    display: grid;
    grid-template-columns: repeat(3, 1fr);
    gap: 1.25rem;
  }
  .card { background: #fff; border: 1px solid #d8d2bd; border-radius: 10px; padding: 1.1rem 1.2rem; }
  .card h3 { margin: 0 0 .2rem; font-size: 1.15rem; }
  .card .when { margin: 0 0 .6rem; font-size: .9rem; color: #5b6650; font-weight: 600; }
  .card p { margin: 0; }
  .table-scroll { overflow-x: auto; max-width: 720px; }
  .fees { width: 100%; min-width: 480px; border-collapse: collapse; background: #fff; }
  .fees th, .fees td { text-align: left; padding: .7rem .9rem; border-bottom: 1px solid #e2dcc7; }
  .fees th { background: #e8e2cc; font-size: .95rem; }
  .note { margin-top: 1rem; font-size: .95rem; color: #4a5642; }
  .contact-box { background: #fff; border: 1px solid #d8d2bd; border-radius: 10px; padding: 1.2rem 1.4rem; }
  .contact-box p { margin: 0 0 .5rem; }
  .email { font-size: 1.1rem; font-weight: 600; overflow-wrap: anywhere; }
  .email a { color: #1f4424; }
  footer { background: #23301f; color: #d8dccc; padding: 1.5rem 0; font-size: .9rem; }

  @media (max-width: 640px) {
    .cards { grid-template-columns: 1fr; }
    .hero h1 { font-size: 2rem; }
  }
</style>
</head>
<body>
<header class="site-header">
  <div class="wrap">
    <span class="brand">Thistlebank Allotments</span>
    <nav>
      <a href="#about">About</a>
      <a href="#events">Events</a>
      <a href="#fees">Plot fees</a>
      <a href="#contact">Contact</a>
    </nav>
  </div>
</header>

<main>
  <div class="hero">
    <div class="wrap">
      <h1>Grow something together</h1>
      <p>A friendly community garden on the edge of town, with forty-two plots, a shared tool shed and a kettle that is always on.</p>
    </div>
  </div>

  <section class="block" id="about">
    <div class="wrap">
      <h2>About the garden</h2>
      <p>Thistlebank has been run by its members since the garden opened. Plots are small enough to manage in an evening a week, and there is a shared bed for anyone who is not ready to commit to a plot of their own.</p>
      <p>Beginners are welcome. Long-time growers are usually happy to explain why your beans are doing that.</p>
    </div>
  </section>

  <section class="block" id="events">
    <div class="wrap">
      <h2>Coming up</h2>
      <div class="cards">
        <article class="card">
          <h3>Seed swap</h3>
          <p class="when">First Saturday of the month, 10:00</p>
          <p>Bring spare seeds, take what you need. Tea and biscuits in the shed.</p>
        </article>
        <article class="card">
          <h3>Compost clinic</h3>
          <p class="when">Second Wednesday, 18:30</p>
          <p>Learn why the heap smells and what to do about it. Bring gloves.</p>
        </article>
        <article class="card">
          <h3>Autumn bake and dig</h3>
          <p class="when">Last Sunday of October, 13:00</p>
          <p>Tidy the shared beds, then eat cake made from whatever we grew.</p>
        </article>
      </div>
    </div>
  </section>

  <section class="block" id="fees">
    <div class="wrap">
      <h2>Plot fees</h2>
      <div class="table-scroll">
        <table class="fees">
          <thead>
            <tr><th>Plot size</th><th>Best for</th><th>Per year</th><th>Water</th></tr>
          </thead>
          <tbody>
            <tr><td>Small (3 x 5 m)</td><td>Salad and herbs</td><td>$35</td><td>Included</td></tr>
            <tr><td>Medium (5 x 8 m)</td><td>A family of four</td><td>$60</td><td>Included</td></tr>
            <tr><td>Large (8 x 10 m)</td><td>Serious growers</td><td>$95</td><td>Included</td></tr>
            <tr><td>Shared bed</td><td>Trying it out</td><td>$10</td><td>Included</td></tr>
          </tbody>
        </table>
      </div>
      <p class="note">Fees are due on the first of March. Concessions are available; just ask.</p>
    </div>
  </section>

  <section class="block" id="contact">
    <div class="wrap">
      <h2>Contact</h2>
      <div class="contact-box">
        <p>Write to the plot secretary about the waiting list, a visit or anything else.</p>
        <p class="email"><a href="mailto:membershipandplotenquiries@thistlebankallotments.example">membershipandplotenquiries@thistlebankallotments.example</a></p>
        <p>Open days are the first Saturday of every month.</p>
      </div>
    </div>
  </section>
</main>

<footer>
  <div class="wrap">Thistlebank Allotments. Run by members, for members.</div>
</footer>
</body>
</html>

Notes:

  • The table is now 100% wide up to 720px, so it looks the same as before on a laptop. On a phone it fits the screen, and if it's still too tight it scrolls within its own box rather than the whole page.
  • I also let the nav links wrap and reduced the hero heading slightly on phones, so the header and hero don't cause overflow on very small screens. These only apply at 640px wide or less (the nav wrapping is harmless on a laptop, where it never triggers).
  • If a heading still sits too close to the header on your phone, increase scroll-margin-top: 6rem in section.block.
  • Floor2/2, 100% passed
  • Middle2/2, 100% passed
  • Top3/3, 100% passed
7 of 7 checks passedfloor 2/2, middle 2/2, top 3/3
  • Passed.
    Returns one whole HTML page that loads
    Floor Page test
  • Passed.
    All the original text is still on the page
    Floor Page test
  • Passed.
    On a 375 px phone the page no longer scrolls sideways
    Middle Page test
  • Passed.
    On a phone the three event cards stack, one per row
    Middle Page test
  • Passed.
    Tapping each top link shows the section heading below the header (phone and laptop)
    Top Page test
  • Passed.
    Laptop layout unchanged: header still sticky, three cards in a row, blocks in the same places
    Top Page test
  • Passed.
    No console errors and nothing loaded from outside
    Top Page test
  • Took 16 s.
  • First words after 1.5 s.
  • Wrote 3,100 tokens.
  • 3.9¢ at API prices.
All numbers.
Time to finish16 s
Time to first words1.5 s
Input tokens2
Cached tokens3,364
Output tokens3,100
Thinking tokens0
Words in the answer931
Price at API rates3.9¢

AI answer from Sonnet 5.5 · max effort · 2 Oct 2026

0 of 7

The answer was cut off at the token limit.

  • Floor0/2, 0% passed
  • Middle0/2, 0% passed
  • Top0/3, 0% passed
0 of 7 checks passedfloor 0/2, middle 0/2, top 0/3
  • Missed.
    Returns one whole HTML page that loadsNo answer
    Floor Page test
  • Missed.
    All the original text is still on the pageNo answer
    Floor Page test
  • Missed.
    On a 375 px phone the page no longer scrolls sidewaysNo answer
    Middle Page test
  • Missed.
    On a phone the three event cards stack, one per rowNo answer
    Middle Page test
  • Missed.
    Tapping each top link shows the section heading below the header (phone and laptop)No answer
    Top Page test
  • Missed.
    Laptop layout unchanged: header still sticky, three cards in a row, blocks in the same placesNo answer
    Top Page test
  • Missed.
    No console errors and nothing loaded from outsideNo answer
    Top Page test
  • Took 4 min 0 s.
  • Wrote 32,000 tokens, 32,000 of them thinking.
  • 33¢ at API prices.
All numbers.
Time to finish4 min 0 s
Time to first wordsnot recorded
Input tokens2
Cached tokens3,364
Output tokens32,000
Thinking tokens32,000
Words in the answer0
Price at API rates33¢

Code

F01Load more shows some posts twice and skips othersLow 7/7Max 7/7Same

What it asks for. Pastes a Python function that pages through a community board's posts, newest first, and reports that some posts show up twice while others never appear, more often on busy days. Many posts share exactly the same timestamp. It asks for the cause and a fixed function with the same name, arguments and return shape, where posts with the same time come higher id first.

Both settings passed 7 of 7 checks.

AI answer from Sonnet 5.5 · low effort · 2 Oct 2026

7 of 7

What was wrong

The cursor only holds created_at, and the query uses created_at <= ? with no tiebreaker. Two things go wrong:

  • Duplicates: the next page starts at the last row's timestamp, inclusive. Every post sharing that timestamp that was already shown comes back again.
  • Missing posts: ORDER BY created_at DESC has no tiebreaker, so rows with equal timestamps can come back in any order. If a tie group straddles a page boundary, the rows on the next page may not be the ones that were skipped. If a whole page shares one timestamp, the cursor can never advance past it. Busy days and batch imports make ties common, which is why it shows up more then.

The fix

Order by (created_at DESC, id DESC), a total order, and make the cursor carry both values. The next page then starts strictly after that position.

import sqlite3

MAX_LIMIT = 100
COLUMNS = "id, author, body, created_at"


def fetch_page(conn, cursor=None, limit=20):
    """Return (rows, next_cursor): one page of the board, newest first.

    rows is a list of (id, author, body, created_at) tuples.
    next_cursor is an opaque string the caller hands back to get the next
    page, or None when there are no more posts.

    Ordering is created_at DESC, then id DESC for posts sharing a created_at.

    posts(id INTEGER PRIMARY KEY, author TEXT, body TEXT, created_at INTEGER)
    created_at is unix seconds.
    """
    limit = max(1, min(int(limit), MAX_LIMIT))

    if cursor is None:
        cur = conn.execute(
            f"SELECT {COLUMNS} FROM posts "
            "ORDER BY created_at DESC, id DESC LIMIT ?",
            (limit + 1,),
        )
    else:
        try:
            created_str, id_str = cursor.split(":", 1)
            last_created, last_id = int(created_str), int(id_str)
        except (ValueError, AttributeError):
            raise ValueError("invalid cursor")
        cur = conn.execute(
            f"SELECT {COLUMNS} FROM posts "
            "WHERE created_at < ? OR (created_at = ? AND id < ?) "
            "ORDER BY created_at DESC, id DESC LIMIT ?",
            (last_created, last_created, last_id, limit + 1),
        )

    rows = cur.fetchall()
    has_more = len(rows) > limit
    rows = rows[:limit]

    next_cursor = f"{rows[-1][3]}:{rows[-1][0]}" if has_more else None
    return rows, next_cursor

Notes

  • The cursor is now a string like "1700000000:42" (created_at and id). The app still just passes it back unchanged. Old cursors (a bare timestamp) will now raise ValueError. If you need to tolerate them during rollout, catch that and restart from page one.
  • This is keyset pagination, so it stays correct when new posts arrive mid-scroll: new posts sort ahead of the cursor and don't shift later pages, which would cause duplicates with OFFSET.
  • For speed on a large table, add an index: CREATE INDEX idx_posts_created_id ON posts(created_at DESC, id DESC);
  • Floor2/2, 100% passed
  • Middle3/3, 100% passed
  • Top2/2, 100% passed
7 of 7 checks passedfloor 2/2, middle 3/3, top 2/2
  • Passed.
    The first page comes back correct, newest first
    Floor Tested
  • Passed.
    Keeps the return shape: string cursor, None at the end
    Floor Tested
  • Passed.
    Paging through everything shows each post exactly once
    Middle Tested
  • Passed.
    Order is newest first and, on ties, higher id first
    Middle Tested
  • Passed.
    Explains that shared created_at values cause the problem
    Middle Tested
  • Passed.
    Still correct at page sizes 1, 2 and 3
    Top Tested
  • Passed.
    New posts added at the top mid-scroll cause no repeats or gaps
    Top Tested
  • Took 7.8 s.
  • First words after 2.0 s.
  • Wrote 1,138 tokens.
  • 1.5¢ at API prices.
All numbers.
Time to finish7.8 s
Time to first words2.0 s
Input tokens2
Cached tokens1,411
Output tokens1,138
Thinking tokens0
Words in the answer431
Price at API rates1.5¢

AI answer from Sonnet 5.5 · max effort · 2 Oct 2026

7 of 7

What was wrong

The cursor was only the last post's created_at. That value isn't unique, since the importer gives many posts the same second, and the query had no tie-breaker.

  • WHERE created_at <= ? re-selects every post with that timestamp, including ones already on the previous page, so you get duplicates. Changing it to < would just turn those into skipped posts.
  • If limit or more posts share a timestamp, the cursor stops advancing because every page ends inside that group. The older posts are never reached.
  • ORDER BY created_at alone leaves the order of equal timestamps undefined, so it isn't guaranteed to be the same from one query to the next.

Busy days and import runs create more same-second posts, which is why it shows up more then.

The fix is to order by the unique pair (created_at DESC, id DESC), put both values in the cursor, and fetch rows strictly after that pair. id alone can't be the cursor, since ids don't follow created_at.

Fixed function

import sqlite3

MAX_LIMIT = 100
COLUMNS = "id, author, body, created_at"


def fetch_page(conn, cursor=None, limit=20):
    """Return (rows, next_cursor): one page of the board, newest first.

    rows is a list of (id, author, body, created_at) tuples, ordered by
    created_at descending; posts with the same created_at come higher id
    first.
    next_cursor is an opaque string the caller hands back to get the next
    page, or None when there are no more posts.

    posts(id INTEGER PRIMARY KEY, author TEXT, body TEXT, created_at INTEGER)
    created_at is unix seconds.
    """
    limit = max(1, min(int(limit), MAX_LIMIT))

    if cursor is None:
        cur = conn.execute(
            f"SELECT {COLUMNS} FROM posts "
            "ORDER BY created_at DESC, id DESC LIMIT ?",
            (limit + 1,),
        )
    else:
        # The cursor is "<created_at>:<id>" of the last row of the previous page.
        try:
            after_created_at, after_id = map(int, cursor.split(":"))
        except ValueError:
            raise ValueError(f"invalid cursor: {cursor!r}") from None
        cur = conn.execute(
            f"SELECT {COLUMNS} FROM posts "
            "WHERE (created_at, id) < (?, ?) "
            "ORDER BY created_at DESC, id DESC LIMIT ?",
            (after_created_at, after_id, limit + 1),
        )

    rows = cur.fetchall()
    has_more = len(rows) > limit
    rows = rows[:limit]

    next_cursor = None
    if has_more:
        last = rows[-1]  # (id, author, body, created_at)
        next_cursor = f"{last[3]}:{last[0]}"
    return rows, next_cursor

Notes

  • The cursor format changed. A cursor from the old code (a bare number) now raises ValueError, so anyone mid-scroll at deploy time has to reload from the top.
  • (created_at, id) < (?, ?) needs SQLite 3.15 or later (2016). On a big table, add CREATE INDEX IF NOT EXISTS idx_posts_created_id ON posts (created_at, id); so SQLite can walk the index in order instead of sorting the whole table for every page.
  • To check the fix, insert about 50 posts with the same created_at and scrambled ids. Page through with limit=10 and confirm every id comes back exactly once.
  • Floor2/2, 100% passed
  • Middle3/3, 100% passed
  • Top2/2, 100% passed
7 of 7 checks passedfloor 2/2, middle 3/3, top 2/2
  • Passed.
    The first page comes back correct, newest first
    Floor Tested
  • Passed.
    Keeps the return shape: string cursor, None at the end
    Floor Tested
  • Passed.
    Paging through everything shows each post exactly once
    Middle Tested
  • Passed.
    Order is newest first and, on ties, higher id first
    Middle Tested
  • Passed.
    Explains that shared created_at values cause the problem
    Middle Tested
  • Passed.
    Still correct at page sizes 1, 2 and 3
    Top Tested
  • Passed.
    New posts added at the top mid-scroll cause no repeats or gaps
    Top Tested
  • Took 3 min 14 s.
  • First words after 3 min 10 s.
  • Wrote 27,199 tokens, 25,976 of them thinking.
  • 28¢ at API prices.
All numbers.
Time to finish3 min 14 s
Time to first words3 min 10 s
Input tokens2
Cached tokens1,411
Output tokens27,199
Thinking tokens25,976
Words in the answer476
Price at API rates28¢
F02The search box shows results for an old queryLow 7/7Max 7/7Same

What it asks for. Pastes a TypeScript controller for a search-as-you-type box with two bugs: typing fast can leave results for an earlier query on screen, and the loading spinner disappears too early. The API client cannot cancel requests. It asks for the whole fixed file with the same exports, the spinner on exactly while the newest search is loading, and code that Node can run with its built-in TypeScript support.

Both settings passed 7 of 7 checks.

AI answer from Sonnet 5.5 · low effort · 2 Oct 2026

7 of 7

Here is the fixed file:

export interface SearchOptions {
  // The API client. It returns a Promise and can't be cancelled.
  fetchResults: (query: string) => Promise<string[]>;
  // Draws the list. Pass an error as the second argument to show the error state.
  render: (results: string[], error?: Error) => void;
  // Shows or hides the spinner.
  setLoading: (loading: boolean) => void;
}

export function createSearch({ fetchResults, render, setLoading }: SearchOptions) {
  // Incremented on every input. A request is only "current" if its id
  // still matches this value when it settles.
  let latestId = 0;

  return {
    // Called on every keystroke with the current text of the box.
    async onInput(raw: string): Promise<void> {
      const id = ++latestId;
      const query = raw.trim();

      if (query === "") {
        // Invalidates any in-flight request, since latestId has moved on.
        setLoading(false);
        render([]);
        return;
      }

      setLoading(true);
      try {
        const results = await fetchResults(query);
        if (id !== latestId) return; // stale: a newer input exists
        render(results);
      } catch (err) {
        if (id !== latestId) return; // stale error: ignore it too
        render([], err instanceof Error ? err : new Error(String(err)));
      } finally {
        if (id === latestId) setLoading(false);
      }
    },
  };
}

What was wrong

Both bugs came from treating every request as if it were the only one.

  1. Stale results. The requests can't be cancelled and may finish in any order. If "ab" was typed after "a" but the "a" response arrived later, it called render last and overwrote the correct list. Each call to onInput now takes a sequence number, and after the await it only renders if it is still the newest one. This applies to errors as well, so an old failure can't replace a newer good list.
  2. Spinner disappearing early. Every request's finally ran setLoading(false), so when an older request finished, it turned the spinner off while the newest was still loading. Now only the newest request can turn it off, and it does so right after rendering its results or error.
  3. Emptied box. The empty-query path also takes a new sequence number, which invalidates anything in flight. It clears the list and turns the spinner off explicitly (it previously didn't touch the spinner, so one started by an earlier keystroke would stay on). Any older request that finishes later is ignored.

The file uses only plain TypeScript syntax that Node's type stripping handles (no enums or other syntax that needs compiling). It still calls fetchResults on every input, and the exported function, its options and onInput are unchanged.

  • Floor2/2, 100% passed
  • Middle2/2, 100% passed
  • Top3/3, 100% passed
7 of 7 checks passedfloor 2/2, middle 2/2, top 3/3
  • Passed.
    The file loads under type stripping and a plain search works
    Floor Tested
  • Passed.
    A failing search still shows the error and hides the spinner
    Floor Tested
  • Passed.
    Only the newest query's results are left on screen
    Middle Tested
  • Passed.
    The spinner stays on until the newest request finishes, then goes off
    Middle Tested
  • Passed.
    An old request that fails does not touch the newer results
    Top Tested
  • Passed.
    Emptying the box clears the list, and late answers don't bring it back
    Top Tested
  • Passed.
    If the newest search fails, the error is shown and stays
    Top Tested
  • Took 7.0 s.
  • First words after 1.4 s.
  • Wrote 912 tokens.
  • 1.3¢ at API prices.
All numbers.
Time to finish7.0 s
Time to first words1.4 s
Input tokens2
Cached tokens1,362
Output tokens912
Thinking tokens0
Words in the answer425
Price at API rates1.3¢

AI answer from Sonnet 5.5 · max effort · 2 Oct 2026

7 of 7

Here is the whole fixed file:

export interface SearchOptions {
  // The API client. It returns a Promise and can't be cancelled.
  fetchResults: (query: string) => Promise<string[]>;
  // Draws the list. Pass an error as the second argument to show the error state.
  render: (results: string[], error?: Error) => void;
  // Shows or hides the spinner.
  setLoading: (loading: boolean) => void;
}

export function createSearch({ fetchResults, render, setLoading }: SearchOptions) {
  // Id of the most recent onInput call. Requests can't be cancelled, so a call
  // may only touch the list or the spinner while its id is still the latest.
  let latestId = 0;

  return {
    // Called on every keystroke with the current text of the box.
    async onInput(raw: string): Promise<void> {
      const query = raw.trim();

      // Every call takes a new id, including one that empties the box, so that
      // all requests started earlier become stale. Do this before any await.
      const id = ++latestId;
      const isLatest = () => id === latestId;

      if (query === "") {
        setLoading(false);
        render([]);
        return;
      }

      setLoading(true);
      try {
        const results = await fetchResults(query);
        if (isLatest()) render(results);
      } catch (err) {
        if (isLatest()) {
          render([], err instanceof Error ? err : new Error(String(err)));
        }
      } finally {
        // Re-checked on purpose: a stale call must not switch off the spinner
        // that now belongs to a newer search.
        if (isLatest()) setLoading(false);
      }
    },
  };
}

What was wrong

  • Stale list: every onInput call awaited its own request and then rendered whatever came back. Responses can arrive in any order, so the last one to land won, not the one for the text currently in the box.
  • Spinner: finally { setLoading(false) } ran for every request. The first old request to finish switched the spinner off while the newest was still loading.
  • Emptying the box: that branch cleared the list but never touched the spinner or the requests already in flight. A slow old response could repaint the list afterwards, and the spinner could stay on until it landed.

The fix

Since the requests can't be cancelled, they run to the end and their outcome is ignored. A counter inside createSearch goes up on every onInput call, including the empty-box one. A call only touches render or setLoading(false) if it is still the latest when it finishes. This covers errors too, so a failed old request shows nothing. Emptying the box makes everything in flight stale and turns the spinner off itself.

I used a counter instead of comparing query strings because the same text can come back (a → ab → a). An old request would then look current while the newest is still loading. The counter is per createSearch instance, so several search boxes won't interfere with each other.

The exports, options and onInput signature are unchanged, and there is still one fetchResults call per non-empty input. The only TypeScript syntax is type annotations, so Node can run the file directly. If you add debouncing later, keep the guard, because debouncing reduces overlapping requests but doesn't make responses arrive in order.

I haven't run this. I traced it by hand for these cases:

  • the newer request finishes first
  • the older request finishes first
  • an older request fails
  • the box is emptied, and then retyped, while a request is in flight
  • Floor2/2, 100% passed
  • Middle2/2, 100% passed
  • Top3/3, 100% passed
7 of 7 checks passedfloor 2/2, middle 2/2, top 3/3
  • Passed.
    The file loads under type stripping and a plain search works
    Floor Tested
  • Passed.
    A failing search still shows the error and hides the spinner
    Floor Tested
  • Passed.
    Only the newest query's results are left on screen
    Middle Tested
  • Passed.
    The spinner stays on until the newest request finishes, then goes off
    Middle Tested
  • Passed.
    An old request that fails does not touch the newer results
    Top Tested
  • Passed.
    Emptying the box clears the list, and late answers don't bring it back
    Top Tested
  • Passed.
    If the newest search fails, the error is shown and stays
    Top Tested
  • Took 2 min 11 s.
  • First words after 2 min 6 s.
  • Wrote 19,021 tokens, 17,868 of them thinking.
  • 19¢ at API prices.
All numbers.
Time to finish2 min 11 s
Time to first words2 min 6 s
Input tokens2
Cached tokens1,362
Output tokens19,021
Thinking tokens17,868
Words in the answer561
Price at API rates19¢
F03Some subscriptions renew a day earlyLow 7/7Max 7/7Same

What it asks for. Pastes a JavaScript function that lists a subscription's monthly renewal dates, with its rules in a comment, and reports two symptoms: customers far from UTC renew a day early, and a customer who signed up on the last day of January has been charged on the 28th of every month since. It asks for the function fixed for Node 22 with no libraries, and a short note on what was wrong.

Both settings passed 7 of 7 checks.

AI answer from Sonnet 5.5 · low effort · 2 Oct 2026

7 of 7

There were two bugs, and the fix needs both the local date and the original day.

// Monthly renewal dates for a subscription.
//
// signupIso  : the moment the customer signed up, as an ISO string in UTC,
//              e.g. "2025-01-31T23:30:00Z"
// timeZone   : the customer's IANA time zone, e.g. "Pacific/Auckland"
// count      : how many upcoming renewals to return
// returns    : an array of count local calendar dates, "YYYY-MM-DD", in the
//              customer's time zone, earliest first
//
// Rules:
//  - The first renewal is one month after the signup, not on the signup day.
//  - A renewal falls on the same day of the month as the signup's local date
//    in the customer's time zone.
//  - If a month is too short for that day, use its last day.
//  - Always count from the original signup day, never from the previous renewal.
export function renewalDates(signupIso, timeZone, count) {
  // Local calendar date of the signup in the customer's time zone.
  const parts = new Intl.DateTimeFormat("en-CA", {
    timeZone,
    year: "numeric",
    month: "2-digit",
    day: "2-digit",
  }).formatToParts(new Date(signupIso));
  const get = (type) => Number(parts.find((p) => p.type === type).value);
  const year = get("year");
  const month = get("month"); // 1-12
  const day = get("day");

  const pad = (n, w = 2) => String(n).padStart(w, "0");
  const dates = [];
  for (let i = 1; i <= count; i++) {
    // Months since year 0, always counted from the original signup month.
    const total = year * 12 + (month - 1) + i;
    const y = Math.floor(total / 12);
    const m = total % 12; // 0-11
    // Day 0 of the next month = last day of this month.
    const lastDay = new Date(Date.UTC(y, m + 1, 0)).getUTCDate();
    dates.push(`${pad(y, 4)}-${pad(m + 1)}-${pad(Math.min(day, lastDay))}`);
  }
  return dates;
}

What was wrong

  1. Wrong time zone. The old code took the day of the month from the UTC date, ignoring timeZone. For the Auckland customer who signed up on the morning of the 1st local time, it is still the 31st in UTC, so they were renewed on the 31st. The fix reads the signup's local calendar date in the customer's time zone using Intl.DateTimeFormat, and builds the output dates from that.
  2. Drift from the previous renewal. The loop read day from d on every pass, and d had already been clamped to a shorter month. A January 31st signup became February 28th, and it stayed on the 28th from then on. The fix takes the day once from the original signup and computes each renewal as signup month + i, clamping to the month's length each time. So a 31st signup gives Feb 28, Mar 31, Apr 30, and so on.

Because the output dates are built from plain year, month and day numbers rather than from a Date in a time zone, daylight saving changes can't shift them.

  • Floor2/2, 100% passed
  • Middle2/2, 100% passed
  • Top3/3, 100% passed
7 of 7 checks passedfloor 2/2, middle 2/2, top 3/3
  • Passed.
    Mid-month signups in UTC get the right dates
    Floor Tested
  • Passed.
    Returns exactly count dates as YYYY-MM-DD, in order
    Floor Tested
  • Passed.
    Uses the customer's local date (Auckland, Kolkata, LA, London)
    Middle Tested
  • Passed.
    Explains both causes: the UTC date and drifting month to month
    Middle Read
  • Passed.
    Signups on the 29th to 31st keep their day after short months
    Top Tested
  • Passed.
    Leap and non-leap Februaries are handled, including 2100
    Top Tested
  • Passed.
    Signups just before or after a clock change get the right local date
    Top Tested
  • Took 7.5 s.
  • First words after 1.4 s.
  • Wrote 1,072 tokens.
  • 1.4¢ at API prices.
All numbers.
Time to finish7.5 s
Time to first words1.4 s
Input tokens2
Cached tokens1,385
Output tokens1,072
Thinking tokens0
Words in the answer485
Price at API rates1.4¢

AI answer from Sonnet 5.5 · max effort · 2 Oct 2026

7 of 7

// Monthly renewal dates for a subscription.
//
// signupIso  : the moment the customer signed up, as an ISO string in UTC,
//              e.g. "2025-01-31T23:30:00Z"
// timeZone   : the customer's IANA time zone, e.g. "Pacific/Auckland"
// count      : how many upcoming renewals to return
// returns    : an array of count local calendar dates, "YYYY-MM-DD", in the
//              customer's time zone, earliest first
//
// Rules:
//  - The first renewal is one month after the signup, not on the signup day.
//  - A renewal falls on the same day of the month as the signup's local date
//    in the customer's time zone.
//  - If a month is too short for that day, use its last day.
//  - Always count from the original signup day, never from the previous renewal.
export function renewalDates(signupIso, timeZone, count) {
  // The signup's calendar date as the customer sees it, not in UTC.
  const parts = new Intl.DateTimeFormat('en-US', {
    timeZone,
    calendar: 'gregory',
    numberingSystem: 'latn',
    year: 'numeric',
    month: 'numeric',
    day: 'numeric',
  }).formatToParts(new Date(signupIso));
  const part = (type) => Number(parts.find((p) => p.type === type).value);
  const signupYear = part('year');
  const signupMonth = part('month') - 1; // 0-11
  const signupDay = part('day');

  const pad = (n, width) => String(n).padStart(width, '0');

  const dates = [];
  for (let i = 1; i <= count; i++) {
    // Always offset from the original signup month, never from the last renewal.
    const monthsFromJan = signupMonth + i;
    const year = signupYear + Math.floor(monthsFromJan / 12);
    const month = monthsFromJan % 12; // 0-11
    const lastDay = new Date(Date.UTC(year, month + 1, 0)).getUTCDate();
    const day = Math.min(signupDay, lastDay);
    dates.push(`${pad(year, 4)}-${pad(month + 1, 2)}-${pad(day, 2)}`);
  }
  return dates;
}

What was wrong

  1. It ignored timeZone. It took the day from the UTC instant (getUTCDate) and returned UTC dates (toISOString). Auckland is 12–13 hours ahead of UTC, so a morning-of-the-1st signup is still the 31st in UTC. That customer was anchored to the 31st and billed a day early. The fix reads the signup's calendar date in the customer's zone with Intl.DateTimeFormat. It then does all the month arithmetic on plain year/month/day numbers, so DST can't shift anything.
  2. It counted from the previous renewal. day was re-read from d on every pass, and d had already been clamped. Once February pushed the 31st down to the 28th, every later month used the 28th. Now the day is read once from the signup, and each renewal is the signup month plus i, clamped to that month's length.

I checked these by hand rather than running them:

renewalDates("2025-01-31T10:00:00Z", "Europe/London", 4);
// ["2025-02-28", "2025-03-31", "2025-04-30", "2025-05-31"]

renewalDates("2025-05-31T20:00:00Z", "Pacific/Auckland", 3);
// local signup is 1 June -> ["2025-07-01", "2025-08-01", "2025-09-01"]

One behaviour change: timeZone is now actually used, so an invalid zone name throws a RangeError instead of being silently ignored. If you have renewal dates already stored from the old function, they will still be wrong until recomputed.

  • Floor2/2, 100% passed
  • Middle2/2, 100% passed
  • Top3/3, 100% passed
7 of 7 checks passedfloor 2/2, middle 2/2, top 3/3
  • Passed.
    Mid-month signups in UTC get the right dates
    Floor Tested
  • Passed.
    Returns exactly count dates as YYYY-MM-DD, in order
    Floor Tested
  • Passed.
    Uses the customer's local date (Auckland, Kolkata, LA, London)
    Middle Tested
  • Passed.
    Explains both causes: the UTC date and drifting month to month
    Middle Read
  • Passed.
    Signups on the 29th to 31st keep their day after short months
    Top Tested
  • Passed.
    Leap and non-leap Februaries are handled, including 2100
    Top Tested
  • Passed.
    Signups just before or after a clock change get the right local date
    Top Tested
  • Took 1 min 19 s.
  • First words after 1 min 13 s.
  • Wrote 12,142 tokens, 10,842 of them thinking.
  • 13¢ at API prices.
All numbers.
Time to finish1 min 19 s
Time to first words1 min 13 s
Input tokens2
Cached tokens1,385
Output tokens12,142
Thinking tokens10,842
Words in the answer491
Price at API rates13¢
F04Review a pull request before mergingLow 6/7Max 7/7Max +1

What it asks for. Pastes a pull request of about 170 lines that adds invoice endpoints to an Express and TypeScript API, and asks for a review before merging that lists what needs fixing in order of importance plus other points worth noting, in under 400 words.

Max effort passed 1 more check than low.

AI answer from Sonnet 5.5 · low effort · 2 Oct 2026

6 of 7

Must fix before merging

  1. Broken access control on GET /invoices/:id. findInvoiceById looks up by id only and never checks account_id. Any logged-in user can read any account's invoice by guessing or enumerating ids. Pass req.user.accountId into the query (WHERE id = $1 AND account_id = $2) and return 404 on a miss, so you don't reveal that the id exists. This is the blocker.
  2. The to date isn't inclusive, contrary to the code comment and docs. The query uses created_at < $3 with to as a bare day, so to=2026-01-31 is treated as midnight at the start of Jan 31 and excludes that whole day. Invoices created on the last day vanish. Fix it by making the upper bound exclusive of the next day: created_at < ($3::date + 1), or compute the next day in code. Also check created_at's type and timezone: if it's timestamptz, the day boundaries depend on the session timezone, so decide whether that's UTC and say so in the docs.
  3. Date validation is regex-only. 2026-99-99 passes DAY and then Postgres throws, producing a 500 rather than a 400. Validate that it's a real calendar date (e.g. parse it and round-trip it). Consider also rejecting from > to.
  4. Async handlers have no error handling. If you're on Express 4, a rejected promise from db.query is never passed to the error middleware and can hang the request or cause an unhandled rejection. Wrap the handlers or use express-async-errors. Express 5 handles this itself. Confirm which version you're on.

Worth knowing

  • The sort whitelist is done correctly: Object.hasOwn guards against __proto__ and similar keys, and column and direction only ever come from fixed values. There's no SQL injection risk there.
  • String(req.query.from ?? "") turns arrays like ?from=a&from=b into odd strings. They fail the regex, so it's safe, but it's slightly sloppy.
  • The list endpoint has no pagination or limit. An account with many invoices returns everything in one response. Add a limit and a cursor or offset before it becomes a problem.
  • Add a tiebreaker to ORDER BY (e.g. , id) so pages and ordering are stable when amounts or dates are equal.
  • The fmtMoney → formatMoney rename is clean, and both callers are updated. If anything outside this diff imports fmtMoney (scripts, tests), it will break, so grep for it.
  • req.user is used without a null check, which is fine because requireLogin runs first, but make sure its typing is declared.
  • There are no tests in this PR. At minimum add tests for cross-account access (item 1) and for the last-day boundary (item 2).
  • Floor1/2, 50% passed
  • Middle2/2, 100% passed
  • Top3/3, 100% passed
6 of 7 checks passedfloor 1/2, middle 2/2, top 3/3
  • Passed.
    Gives a clear verdict: do not merge as it is
    Floor Counted
  • Missed.
    Stays under the 400-word limit420 words
    Floor Counted
  • Passed.
    Spots that any signed-in user can read another account's invoice
    Middle Counted
  • Passed.
    Gives a concrete fix: scope the lookup by account or return 404/403
    Middle Counted
  • Passed.
    Spots that the end date is left out although the docs say inclusive
    Top Counted
  • Passed.
    Puts the invoice access hole first, ahead of the other issues
    Top Read
  • Passed.
    Does not call the ORDER BY string building an injection risk
    Top Read
  • Took 8.3 s.
  • First words after 1.5 s.
  • Wrote 960 tokens.
  • 1.8¢ at API prices.
All numbers.
Time to finish8.3 s
Time to first words1.5 s
Input tokens2
Cached tokens3,194
Output tokens960
Thinking tokens0
Words in the answer437
Price at API rates1.8¢

AI answer from Sonnet 5.5 · max effort · 2 Oct 2026

7 of 7

Fix before merging (most important first)

  1. Any logged-in user can read any account's invoice. findInvoiceById filters on id alone, and GET /invoices/:id never checks the owner. Pass req.user.accountId in and query WHERE id = $1 AND account_id = $2. The existing 404 branch then covers other accounts' invoices without revealing that the ID exists. (The list endpoint is scoped correctly.)
  2. to isn't inclusive, despite the comment and docs. created_at < $3 with to=2026-01-31 excludes everything from 00:00 on Jan 31, so from == to always returns nothing. Use created_at < ($3::date + 1).
  3. Validation accepts impossible dates. The regex lets 2026-02-31 through. The database then throws, giving a 500 (or an unhandled rejection, depending on how your app handles async errors) instead of your bad_date 400. If invoices.id is a uuid column, /:id has the same problem with malformed IDs. Swap the DAY.test calls for a real date check:
const isDay = (s: string) => {
  const d = new Date(`${s}T00:00:00Z`);
  return DAY.test(s) && !Number.isNaN(+d) && d.toISOString().startsWith(s);
};

Worth knowing

  • Check the rename before merging. fmtMoney → formatMoney only updates receipts.ts in this diff. I can't see the rest of the repo, so grep for leftover fmtMoney callers and run tsc.
  • The dynamic ORDER BY is safe. The whitelist plus Object.hasOwn blocks keys like __proto__, and dir maps to a literal. No change needed.
  • There's no pagination or limit, so a wide date range returns every invoice. Add an id tiebreaker to the ORDER BY when you paginate.
  • Day boundaries use the database session's time zone. Confirm that's what you want.
  • There are no tests. A cross-account fetch test would have caught #1, and a from == to test would have caught #2.
  • The docs don't say from and to are required, or list the 400/404 responses.
  • Floor2/2, 100% passed
  • Middle2/2, 100% passed
  • Top3/3, 100% passed
7 of 7 checks passedfloor 2/2, middle 2/2, top 3/3
  • Passed.
    Gives a clear verdict: do not merge as it is
    Floor Counted
  • Passed.
    Stays under the 400-word limit
    Floor Counted
  • Passed.
    Spots that any signed-in user can read another account's invoice
    Middle Counted
  • Passed.
    Gives a concrete fix: scope the lookup by account or return 404/403
    Middle Counted
  • Passed.
    Spots that the end date is left out although the docs say inclusive
    Top Counted
  • Passed.
    Puts the invoice access hole first, ahead of the other issues
    Top Read
  • Passed.
    Does not call the ORDER BY string building an injection risk
    Top Read
  • Took 2 min 48 s.
  • First words after 2 min 46 s.
  • Wrote 25,235 tokens, 24,466 of them thinking.
  • 26¢ at API prices.
All numbers.
Time to finish2 min 48 s
Time to first words2 min 46 s
Input tokens2
Cached tokens3,194
Output tokens25,235
Thinking tokens24,466
Words in the answer308
Price at API rates26¢
F05SQL: which customers left and came backLow 7/7Max 7/7Same

What it asks for. Gives the schema of a subscriptions database and asks for one SQLite query that lists each gap in a customer's access that ended with them returning, with the last day of access, the day it came back and the number of days without it. Plan changes can make periods touch, overlap or sit inside one another, and none of that counts as a gap. A small example with its expected output is included.

Both settings passed 7 of 7 checks.

AI answer from Sonnet 5.5 · low effort · 2 Oct 2026

7 of 7

WITH ordered AS (
  SELECT
    customer_id,
    start_date,
    MAX(COALESCE(end_date, '9999-12-31')) OVER (
      PARTITION BY customer_id
      ORDER BY start_date, id
      ROWS BETWEEN UNBOUNDED PRECEDING AND 1 PRECEDING
    ) AS prev_end
  FROM subscription_periods
)
SELECT
  c.name                                                      AS name,
  o.prev_end                                                  AS last_access,
  o.start_date                                                AS came_back,
  CAST(julianday(o.start_date) - julianday(o.prev_end) - 1 AS INTEGER) AS days_without
FROM ordered AS o
JOIN customers AS c ON c.id = o.customer_id
WHERE o.prev_end IS NOT NULL
  AND o.start_date > date(o.prev_end, '+1 day')
ORDER BY c.name, o.start_date, c.id;

How it works

  • For each period, the window function looks at all the customer's earlier periods (ordered by start date, with id only as a tie-breaker) and takes the latest end date among them. This is prev_end, the last day of access the customer had accumulated before this period began.
  • Using the running maximum rather than just the previous row's end date is what handles overlaps and periods nested inside longer ones. A short period inside a longer one can't pull prev_end backwards.
  • A NULL end date (still active) is treated as 9999-12-31. Once a customer has an open-ended period, no later period can start after prev_end, so no gap is reported. A customer who left and never returned has no later period, so no row is produced.
  • A gap exists only when a period starts later than the day after prev_end. Touching periods (start = end + 1), same-day starts and overlaps all fail that test, so they aren't counted.
  • Each gap is reported once, at the first period that starts after it. last_access is prev_end, came_back is that period's start date, and days_without is the day difference minus 1 (e.g. 2024-06-30 → 2024-07-10 gives 10 − 1 = 9).
  • Sorting is by name, then came_back, then customer id, using SQLite's default binary collation. That puts all uppercase letters before lowercase ones, so if names have mixed case and you want a case-insensitive order, use ORDER BY c.name COLLATE NOCASE, ....

On your sample data this returns the three expected rows: Amara 9 days, and Cleo 1 day and 31 days. Bruno is excluded because his periods overlap and the latest one is still active.

  • Floor2/2, 100% passed
  • Middle2/2, 100% passed
  • Top3/3, 100% passed
7 of 7 checks passedfloor 2/2, middle 2/2, top 3/3
  • Passed.
    The query runs and returns the four named columns
    Floor Tested
  • Passed.
    It gets the example in the question right
    Floor Tested
  • Passed.
    Periods that touch or overlap are not counted as gaps
    Middle Tested
  • Passed.
    Several gaps per customer, same-name customers, sort order
    Middle Tested
  • Passed.
    A period inside a longer one does not hide or invent a gap
    Top Tested
  • Passed.
    Open-ended (NULL end) periods never produce a gap after them
    Top Tested
  • Passed.
    Days without access is exact across month, leap and year ends
    Top Tested
  • Took 11 s.
  • First words after 6.5 s.
  • Wrote 1,446 tokens, 552 of them thinking.
  • 1.9¢ at API prices.
All numbers.
Time to finish11 s
Time to first words6.5 s
Input tokens2
Cached tokens1,685
Output tokens1,446
Thinking tokens552
Words in the answer362
Price at API rates1.9¢

AI answer from Sonnet 5.5 · max effort · 2 Oct 2026

7 of 7

WITH running AS (
  SELECT
    customer_id,
    start_date,
    -- latest end date among this customer's periods that sort before this one
    -- (NULL end_date = still active, treated as open-ended)
    MAX(COALESCE(end_date, '9999-12-31')) OVER (
      PARTITION BY customer_id
      ORDER BY start_date
      ROWS BETWEEN UNBOUNDED PRECEDING AND 1 PRECEDING
    ) AS prev_end
  FROM subscription_periods
),
gaps AS (
  SELECT
    customer_id,
    prev_end   AS last_access,
    start_date AS came_back,
    CAST(julianday(start_date) - julianday(prev_end) AS INTEGER) - 1 AS days_without
  FROM running
)
SELECT
  c.name         AS name,
  g.last_access  AS last_access,
  g.came_back    AS came_back,
  g.days_without AS days_without
FROM gaps AS g
JOIN customers AS c ON c.id = g.customer_id
WHERE g.days_without > 0
ORDER BY c.name COLLATE NOCASE, g.came_back, c.id;

How it works

  • Running maximum end date: running gives each period a prev_end, the latest end date among that customer's periods that start before it. The frame stops one row short of the current row. A running maximum, rather than the previous row's end date, is what makes touching, same-day, overlapping and fully nested periods harmless. It never looks at ids. A NULL end date counts as 9999-12-31, so nothing can follow an active period as a gap.
  • Counting whole days: gaps counts the days between prev_end and the period's start, minus 1. Touching periods give 0, and overlaps or same-day starts give negative numbers. Only positive values survive WHERE g.days_without > 0. A customer's first period has no prev_end (NULL), so it drops out there too.
  • What each row means: each surviving row is the first period after a gap. came_back is its start date and last_access is the prev_end before it. A customer who left and never returned has no later period, so they produce no row.
  • Sorting: names sort case-insensitively (COLLATE NOCASE), so "de Souza" lands under D instead of after every capitalised name. Delete that clause if you want strict byte order. Ties then go by came_back and customer id.

I traced your sample by hand and it gives the three expected rows (Amara 9, Cleo 1, Cleo 31). It assumes no period has an end date earlier than its start date. It uses only window functions (SQLite 3.25+) and the built-in date functions, so 3.53 is fine.

  • Floor2/2, 100% passed
  • Middle2/2, 100% passed
  • Top3/3, 100% passed
7 of 7 checks passedfloor 2/2, middle 2/2, top 3/3
  • Passed.
    The query runs and returns the four named columns
    Floor Tested
  • Passed.
    It gets the example in the question right
    Floor Tested
  • Passed.
    Periods that touch or overlap are not counted as gaps
    Middle Tested
  • Passed.
    Several gaps per customer, same-name customers, sort order
    Middle Tested
  • Passed.
    A period inside a longer one does not hide or invent a gap
    Top Tested
  • Passed.
    Open-ended (NULL end) periods never produce a gap after them
    Top Tested
  • Passed.
    Days without access is exact across month, leap and year ends
    Top Tested
  • Took 3 min 41 s.
  • First words after 3 min 37 s.
  • Wrote 28,998 tokens, 28,071 of them thinking.
  • 29¢ at API prices.
All numbers.
Time to finish3 min 41 s
Time to first words3 min 37 s
Input tokens2
Cached tokens1,685
Output tokens28,998
Thinking tokens28,071
Words in the answer369
Price at API rates29¢
F06A contacts import that survives real CSV filesLow 9/9Max 0/9Max −9

What it asks for. Asks for a Python function, standard library only, that imports contacts from CSV files exported by many different tools and returns each contact's name, email and phone. The files vary in header names, comma or semicolon separators, a byte order mark, quoted fields with line breaks, Windows line endings and older files in Windows-1252. Rows without an email are skipped, and a repeated email keeps its first row.

Max effort passed 9 fewer checks than low.

AI answer from Sonnet 5.5 · low effort · 2 Oct 2026

9 of 9

import csv
import io

_NAME_HEADERS = {"name", "full name", "contact name"}
_EMAIL_HEADERS = {"email", "e-mail", "email address", "e-mail address"}
_PHONE_HEADERS = {"phone", "phone number", "mobile", "mobile phone", "telephone"}


def _norm(s):
    return " ".join(s.split()).lower()


def import_contacts(path):
    with open(path, "rb") as f:
        raw = f.read()

    try:
        text = raw.decode("utf-8-sig")
    except UnicodeDecodeError:
        # Windows-1252 has a few undefined bytes; map them through as latin-1
        try:
            text = raw.decode("cp1252")
        except UnicodeDecodeError:
            text = raw.decode("latin-1")
        if text.startswith("\ufeff"):
            text = text[1:]

    if not text.strip():
        return []

    # Header line decides the separator (first line, ignoring quoted content)
    header_line = text.split("\n", 1)[0]
    in_q = False
    commas = semis = 0
    for ch in header_line:
        if ch == '"':
            in_q = not in_q
        elif not in_q:
            if ch == ",":
                commas += 1
            elif ch == ";":
                semis += 1
    delim = ";" if semis > commas else ","

    reader = csv.reader(io.StringIO(text, newline=""), delimiter=delim, quotechar='"')
    rows = iter(reader)

    header = None
    for row in rows:
        if any(c.strip() for c in row):
            header = row
            break
    if header is None:
        return []

    idx = {"name": None, "email": None, "phone": None}
    for i, h in enumerate(header):
        n = _norm(h)
        if n in _NAME_HEADERS and idx["name"] is None:
            idx["name"] = i
        elif n in _EMAIL_HEADERS and idx["email"] is None:
            idx["email"] = i
        elif n in _PHONE_HEADERS and idx["phone"] is None:
            idx["phone"] = i

    if idx["email"] is None:
        return []

    def cell(row, key):
        i = idx[key]
        if i is None or i >= len(row):
            return ""
        return row[i].replace("\r\n", "\n").replace("\r", "\n").strip()

    result = []
    seen = set()
    for row in rows:
        email = cell(row, "email").lower()
        if not email or email in seen:
            continue
        seen.add(email)
        result.append({
            "name": cell(row, "name"),
            "email": email,
            "phone": cell(row, "phone"),
        })
    return result

How it works: The file is read as bytes and decoded as UTF-8 (dropping any BOM), falling back to Windows-1252. The separator is chosen by counting unquoted commas vs. semicolons in the header line. csv.reader then handles quoting, embedded separators, doubled quotes and line breaks (normalized to \n). Columns are found by case- and space-insensitive header matching, and rows are stripped, skipped if they have no email, de-duplicated on the lower-cased email, and kept in file order. A file without an email column returns an empty list.

  • Floor2/2, 100% passed
  • Middle3/3, 100% passed
  • Top4/4, 100% passed
9 of 9 checks passedfloor 2/2, middle 3/3, top 4/4
  • Passed.
    A plain comma-separated file is imported
    Floor Tested
  • Passed.
    Fields are stripped, emails lower-cased, rows without an email skipped
    Floor Tested
  • Passed.
    Header names vary in case, spacing, wording and order
    Middle Tested
  • Passed.
    A UTF-8 BOM and Windows line endings are handled
    Middle Tested
  • Passed.
    Quoted commas and doubled quotes are read correctly
    Middle Tested
  • Passed.
    Line breaks inside quoted fields are kept as one \n
    Top Tested
  • Passed.
    Semicolon files are read, even with commas in fields
    Top Tested
  • Passed.
    Windows-1252 files decode correctly and UTF-8 stays intact
    Top Tested
  • Passed.
    Duplicates, blank and short rows, and a file with every quirk at once
    Top Tested
  • Took 7.1 s.
  • First words after 2.0 s.
  • Wrote 1,108 tokens.
  • 1.5¢ at API prices.
All numbers.
Time to finish7.1 s
Time to first words2.0 s
Input tokens2
Cached tokens1,397
Output tokens1,108
Thinking tokens0
Words in the answer372
Price at API rates1.5¢

AI answer from Sonnet 5.5 · max effort · 2 Oct 2026

0 of 9

The answer was cut off at the token limit.

  • Floor0/2, 0% passed
  • Middle0/3, 0% passed
  • Top0/4, 0% passed
0 of 9 checks passedfloor 0/2, middle 0/3, top 0/4
  • Missed.
    A plain comma-separated file is importedNo answer
    Floor Tested
  • Missed.
    Fields are stripped, emails lower-cased, rows without an email skippedNo answer
    Floor Tested
  • Missed.
    Header names vary in case, spacing, wording and orderNo answer
    Middle Tested
  • Missed.
    A UTF-8 BOM and Windows line endings are handledNo answer
    Middle Tested
  • Missed.
    Quoted commas and doubled quotes are read correctlyNo answer
    Middle Tested
  • Missed.
    Line breaks inside quoted fields are kept as one \nNo answer
    Top Tested
  • Missed.
    Semicolon files are read, even with commas in fieldsNo answer
    Top Tested
  • Missed.
    Windows-1252 files decode correctly and UTF-8 stays intactNo answer
    Top Tested
  • Missed.
    Duplicates, blank and short rows, and a file with every quirk at onceNo answer
    Top Tested
  • Took 3 min 44 s.
  • Wrote 32,000 tokens, 32,000 of them thinking.
  • 32¢ at API prices.
All numbers.
Time to finish3 min 44 s
Time to first wordsnot recorded
Input tokens2
Cached tokens1,397
Output tokens32,000
Thinking tokens32,000
Words in the answer0
Price at API rates32¢
F07Make a slow script fastLow 7/7Max 0/7Max −7

What it asks for. Pastes a Python function that matches bank payments to open invoices by customer, amount and a due date within 7 days, and says it is far too slow on about 200,000 of each. It asks for a replacement that runs in a couple of seconds, returns exactly the same results for every input and leaves the lists it is given unchanged.

Max effort passed 7 fewer checks than low.

AI answer from Sonnet 5.5 · low effort · 2 Oct 2026

7 of 7

from bisect import bisect_left
from datetime import timedelta


def match_payments(payments, invoices):
    """Match bank payments to open invoices.

    payments: list of dicts with id, customer, amount (integer cents) and date (datetime.date).
    invoices: list of dicts with id, customer, amount (integer cents) and due (datetime.date).

    Payments are handled one at a time, in the order they are given. A payment is
    matched to the unpaid invoice that has the same customer and the same amount,
    whose due date is within 7 days either side of the payment date (7 days
    apart still counts). If several invoices qualify, the one with the earliest
    due date wins; if they are due on the same day, the lowest invoice id wins.
    An invoice can be matched to one payment only.

    Returns (matches, unmatched):
      matches   - list of (payment_id, invoice_id), in payment order
      unmatched - list of ids of the payments that found no invoice, in payment order
    """
    week = timedelta(days=7)

    # Group invoices by (customer, amount); sort each group by (due, id).
    raw = {}
    for inv in invoices:
        raw.setdefault((inv["customer"], inv["amount"]), []).append(inv)

    groups = {}
    for key, lst in raw.items():
        lst.sort(key=lambda v: (v["due"], v["id"]))  # stable, as the original's first-seen tie behaviour
        n = len(lst)
        groups[key] = (
            [v["due"] for v in lst],   # sorted due dates
            [v["id"] for v in lst],    # ids in the same order
            list(range(n + 1)),        # "next candidate" links (union-find), sentinel at n
            n,
        )

    used = set()
    matches = []
    unmatched = []

    for pay in payments:
        g = groups.get((pay["customer"], pay["amount"]))
        if g is None:
            unmatched.append(pay["id"])
            continue
        dues, ids, parent, n = g
        pdate = pay["date"]

        start = bisect_left(dues, pdate - week)

        # Find the first index >= start whose invoice is not used.
        r = start
        while r < n:
            if parent[r] != r:
                r = parent[r]
            elif ids[r] in used:
                parent[r] = r + 1
                r += 1
            else:
                break
        # Path compression.
        j = start
        while j < r:
            nj = parent[j]
            parent[j] = r
            j = nj

        if r >= n or dues[r] > pdate + week:
            unmatched.append(pay["id"])
        else:
            used.add(ids[r])
            matches.append((pay["id"], ids[r]))

    return matches, unmatched

Why the old one was slow: for every payment it scanned the whole invoice list and ran the checks on each invoice. That is about 200,000 × 200,000 = 40 billion iterations, so it never finishes in a reasonable time.

What I changed:

  • Grouping: invoices are grouped once by (customer, amount), so a payment only looks at invoices that could match on those two fields.
  • Sorting: each group is sorted once by (due, id). The first usable invoice in that order is exactly the one the old tie-break rules would pick (earliest due date, then lowest id).
  • Window lookup: bisect jumps straight to the first invoice due on or after payment date − 7 days. If that invoice is due after payment date + 7 days, nothing qualifies and the payment is unmatched.
  • Used invoices: a large customer can have thousands of same-amount invoices, so skipping used ones one by one could become slow again. Each group keeps a "next unused" link with path compression, so used invoices are skipped in near-constant time. Used invoices are still tracked by id, as before.

The total cost is roughly O((P + I) log I), which should take a second or two for 200,000 of each.

Behaviour preserved:

  • Payments are processed in the given order.
  • The 7-day limit is inclusive.
  • An invoice is used at most once.
  • The inputs are not modified. Only new lists are built, and the sorting is done on a copy.
  • Floor2/2, 100% passed
  • Middle3/3, 100% passed
  • Top2/2, 100% passed
7 of 7 checks passedfloor 2/2, middle 3/3, top 2/2
  • Passed.
    Same results as the original on simple cases
    Floor Tested
  • Passed.
    Same results as the original on many random small inputs
    Floor Tested
  • Passed.
    The 7-day limit is identical at both edges
    Middle Tested
  • Passed.
    The two input lists are left unchanged
    Middle Tested
  • Passed.
    Fast on 20,000 payments and invoices, same results
    Middle Tested
  • Passed.
    Earliest due date, then lowest id, in payment order, each invoice once
    Top Tested
  • Passed.
    Under 6 seconds on 200,000 payments and invoices, same results
    Top Tested
  • Took 21 s.
  • First words after 13 s.
  • Wrote 2,609 tokens, 1,202 of them thinking.
  • 3.0¢ at API prices.
All numbers.
Time to finish21 s
Time to first words13 s
Input tokens2
Cached tokens1,564
Output tokens2,609
Thinking tokens1,202
Words in the answer604
Price at API rates3.0¢

AI answer from Sonnet 5.5 · max effort · 2 Oct 2026

0 of 7

The answer was cut off at the token limit.

  • Floor0/2, 0% passed
  • Middle0/3, 0% passed
  • Top0/2, 0% passed
0 of 7 checks passedfloor 0/2, middle 0/3, top 0/2
  • Missed.
    Same results as the original on simple casesNo answer
    Floor Tested
  • Missed.
    Same results as the original on many random small inputsNo answer
    Floor Tested
  • Missed.
    The 7-day limit is identical at both edgesNo answer
    Middle Tested
  • Missed.
    The two input lists are left unchangedNo answer
    Middle Tested
  • Missed.
    Fast on 20,000 payments and invoices, same resultsNo answer
    Middle Tested
  • Missed.
    Earliest due date, then lowest id, in payment order, each invoice onceNo answer
    Top Tested
  • Missed.
    Under 6 seconds on 200,000 payments and invoices, same resultsNo answer
    Top Tested
  • Took 4 min 1 s.
  • Wrote 32,000 tokens, 32,000 of them thinking.
  • 32¢ at API prices.
All numbers.
Time to finish4 min 1 s
Time to first wordsnot recorded
Input tokens2
Cached tokens1,564
Output tokens32,000
Thinking tokens32,000
Words in the answer0
Price at API rates32¢

Planning and architecture

A01Plan a booking app for a dog groomerLow 7/9Max 0/9Max −7

What it asks for. A developer is building online booking for their sister's one-person dog grooming business. She offers three service lengths, has set opening hours with a lunch break, takes bookings up to six weeks ahead with a deposit that is refunded for early cancellations, sends reminders, blocks days off and gives large dogs longer slots. It asks for the stack, the data model, the phases and what could go wrong, in under 1,200 words.

Max effort passed 7 fewer checks than low.

AI answer from Sonnet 5.5 · low effort · 2 Oct 2026

7 of 9

Dog-Grooming Booking System: Plan

Guiding principle

Build the smallest thing that takes real bookings and real money, then add comfort. A one-person business doesn't need multi-staff, multi-location or accounts. Avoid generalizing.

Stack

Pick boring, hosted pieces so you aren't running servers on weekends.

  • Backend: Python with FastAPI (or Django, if you want the admin and auth for free). Django is a good fit here: its admin gives you a rough owner dashboard on day one.
  • Database: PostgreSQL (managed: Neon, Supabase or Render). SQLite is tempting, but you want transactional guarantees and easy backups.
  • Frontend: A mobile-first web app, with no native app. Server-rendered pages with a bit of HTMX or Alpine.js are enough. If you prefer, use React/Next.js, but it's more moving parts. Make it a PWA (installable to her home screen) so the owner view feels like an app.
  • Payments: Stripe Checkout (hosted page), with webhooks to confirm. You never touch card data.
  • Reminders: SMS via Twilio, or email via Postmark/Resend. SMS is more reliable for reminders, but costs money and has registration hurdles. Start with email, add SMS later. Run it from a scheduled job (cron on the host, or a worker like APScheduler).
  • Hosting: Render, Fly.io or Railway. One web service, one Postgres, one cron job.
  • Auth: None for clients (booking by email plus a signed link to manage or cancel). Password or magic-link login for your sister only.

Data model

Store every instant in UTC, and keep the business time zone (e.g. Europe/London) as config. Convert only at the edges.

  • Service: id, name, base_duration_min (30/60/90), price_cents, active
  • Client: id, name, email, phone
  • Dog: id, client_id, name, size (small_medium | large), notes
  • Booking: id, client_id, dog_id, service_id, start_at (UTC), end_at (UTC), status (pending_payment | confirmed | cancelled_refunded | cancelled_kept | completed | no_show), duration_min (computed and stored), price_cents, deposit_cents, stripe_payment_intent_id, manage_token, reminder_sent_at, created_at, expires_at (for pending holds)
  • WeeklyHours: weekday, open_time, close_time (local time). Tuesday to Saturday: 09:00 to 17:00.
  • Break: weekday, start_time, end_time (lunch 12:30 to 13:30). Could be folded into WeeklyHours as multiple windows per day.
  • TimeOff (blocked periods): start_date, end_date (or datetime range), reason
  • Settings: timezone, booking_horizon_days (42), cancellation_window_hours (24), deposit_percent (20), slot_granularity_min (e.g. 30)

Key design choices:

  • Store duration and price on the booking. If she later changes a service or the large-dog rule, existing bookings stay correct.
  • Large dog = 2x duration. A 90-minute service becomes 180 minutes. Check that this actually fits: it can't straddle lunch (unless she says it can), and it can't run past 17:00. Large dogs on a 90-minute service are only bookable in 9:00-12:30 (starting by 9:30) or 13:30-17:00 (starting by 14:00). That's a real constraint, so confirm it with her. Also ask whether the price doubles.
  • Prevent double booking in the database, not only in code. In Postgres, use an exclusion constraint on a tstzrange(start_at, end_at) for active statuses (needs the btree_gist extension). This makes races impossible.

Availability logic

Compute slots on demand, not stored:

  1. For the requested date, take the local working windows (open hours minus the lunch break).
  2. Subtract TimeOff and existing active bookings (including unexpired pending_payment holds).
  3. For the needed duration (after the large-dog multiplier), list start times on the slot grid where the whole interval fits within a working window.
  4. Reject dates in the past or beyond 42 days.

Build this as one pure function with lots of unit tests. It is the heart of the system.

Payment and cancellation flow

  1. Client picks service, dog size and slot, then enters details.
  2. Create a pending_payment booking holding the slot for ~10 minutes, then start Stripe Checkout for the 20% deposit.
  3. The Stripe webhook (checkout.session.completed) flips it to confirmed. Never trust the browser redirect alone.
  4. Expired holds are released by a cleanup job (or simply ignored by the availability query because of expires_at).
  5. Cancellation via the manage link: if more than 24 hours before start_at, issue a Stripe refund and set cancelled_refunded. Otherwise set cancelled_kept. The server decides, using the clock, not the client.
  6. Her own cancellations (sick day) should always refund, so give her a separate action.

Deposit on 20% of what? Decide whether it's of the doubled price for large dogs, and round consistently (cents).

Phases

Phase 1: Availability engine (weekend 1-2). Models, hours, lunch, time off, durations, large-dog rule, slot calculation, with tests, including DST cases. No UI yet beyond a plain page listing slots.

Phase 2: Booking without money (weekend 3). Booking form, DB constraint against double booking, confirmation email, manage link. Test on yourself.

Phase 3: Payments (weekend 4-5). Stripe Checkout, webhooks, pending holds, expiry, refunds on cancellation. Use Stripe test mode throughout. Make webhook handling idempotent.

Phase 4: Her side (weekend 6). Mobile login, today/week view, block days, cancel a booking, mark no-show or completed. Test on her actual phone.

Phase 5: Reminders and polish (weekend 7). Day-before reminder job, error alerting, backups, terms/cancellation policy text, privacy notice.

Launch: a soft launch with a few friendly clients while she keeps her current system as a backup for a couple of weeks.

Later: SMS, Google Calendar sync, waitlist, rebooking in one tap.

What could go wrong

Daylight saving time (the big one).

  • Store UTC, define hours in local wall-clock time, and generate slots by building local datetimes with a proper tz library (zoneinfo) and then converting. Never add 24 hours to a timestamp to get "tomorrow".
  • Since she's open Tuesday to Saturday, the clock changes on Sunday and won't touch her working hours directly, but a booking made before the change for a day after it must still land at the right local time. Also, "24 hours before" for refunds and "the day before" for reminders are computed in absolute time, so a booking at 9:00 Tuesday after a Sunday change has a slightly different gap than you'd expect. Test both changeovers.
  • Display times in the business time zone, not the visitor's browser zone.

Race conditions. Two clients grab the same slot, or the webhook arrives after the hold expires. The DB constraint covers the first. For the second, if payment succeeds on an expired hold, either re-check the slot and confirm, or auto-refund.

Webhooks. Stripe retries and delivers out of order. Verify signatures, and make handlers idempotent.

Reminder reliability. Cron jobs fail silently. Record reminder_sent_at, run the job often (hourly) selecting "due and not yet sent", and alert on failures. Avoid double sends. Define "the day before" explicitly (e.g. 10:00 local the previous day), and think about Tuesday bookings: the day before is Monday, when she's closed, but the reminder still goes out.

Large-dog ambiguity. Clients self-report size and may get it wrong. Add a note ("we'll confirm size at drop-off") and let her edit a booking's size or duration. Decide what happens if the dog turns out to be large, with no free time after.

Policy edge cases. Rescheduling (is it cancel and rebook?), no-shows, a late cancellation caused by her blocking a day that already has bookings. When she blocks a day, the system should list the affected bookings and prompt refund and notification, rather than silently leaving them.

Money and law. Refund fees (Stripe typically doesn't return its processing fee, so she may lose a little on each refund). Consumer cancellation rules vary by country, so write the policy clearly and show it before payment. Store minimal personal data and say so.

Phone usability. Most of the effort for her is the owner screen. Keep it to a handful of big-button actions, and test with her early rather than at the end.

Scope creep and maintenance. You'll be her IT support. Keep dependencies few, set up automatic backups, uptime monitoring and error tracking (Sentry free tier), and write a one-page "if it breaks" note.

Questions to ask her before you start

  1. Does the large-dog rule double the price as well as the time?
  2. Can appointments span the lunch break, or must they fit around it?
  3. Does she want a buffer between appointments (cleaning, breaks)?
  4. Does she want to approve bookings, or are they instant?
  5. SMS or email reminders?
  6. What's her current policy on late arrivals and no-shows?

Answering these before Phase 1 will save you rework, particularly the first three, since they change the availability logic.

  • Floor2/2, 100% passed
  • Middle2/4, 50% passed
  • Top3/3, 100% passed
7 of 9 checks passedfloor 2/2, middle 2/4, top 3/3
  • Passed.
    Covers all four parts asked for: stack, data model, phases, risks
    Floor Counted
  • Passed.
    The data model has bookings, services and clients or dogs
    Floor Counted
  • Missed.
    Stays under the 1,200-word limit1369 words
    Middle Counted
  • Passed.
    Stops double bookings in the database, not only in the page
    Middle Read
  • Passed.
    Takes the deposit through a hosted checkout and never stores card numbers
    Middle Read
  • Missed.
    The first phase is small and usable before payments existPhase 1 is only the availability engine, with "No UI yet beyond a plain page listing slots". Booking itself waits until…
    Middle Read
  • Passed.
    Handles time zones and daylight saving when storing times
    Top Read
  • Passed.
    Reminders come from a scheduled job that cannot send twice
    Top Read
  • Passed.
    Applies the 24-hour refund rule from the appointment start time
    Top Read
  • Took 29 s.
  • First words after 1.6 s.
  • Wrote 3,283 tokens.
  • 3.5¢ at API prices.
All numbers.
Time to finish29 s
Time to first words1.6 s
Input tokens2
Cached tokens1,048
Output tokens3,283
Thinking tokens0
Words in the answer1,406
Price at API rates3.5¢

AI answer from Sonnet 5.5 · max effort · 2 Oct 2026

0 of 9

The answer was cut off at the token limit.

  • Floor0/2, 0% passed
  • Middle0/4, 0% passed
  • Top0/3, 0% passed
0 of 9 checks passedfloor 0/2, middle 0/4, top 0/3
  • Missed.
    Covers all four parts asked for: stack, data model, phases, risksNo answer
    Floor Counted
  • Missed.
    The data model has bookings, services and clients or dogsNo answer
    Floor Counted
  • Missed.
    Stays under the 1,200-word limitNo answer
    Middle Counted
  • Missed.
    Stops double bookings in the database, not only in the pageNo answer
    Middle Read
  • Missed.
    Takes the deposit through a hosted checkout and never stores card numbersNo answer
    Middle Read
  • Missed.
    The first phase is small and usable before payments existNo answer
    Middle Read
  • Missed.
    Handles time zones and daylight saving when storing timesNo answer
    Top Read
  • Missed.
    Reminders come from a scheduled job that cannot send twiceNo answer
    Top Read
  • Missed.
    Applies the 24-hour refund rule from the appointment start timeNo answer
    Top Read
  • Took 4 min 0 s.
  • Wrote 32,000 tokens, 32,000 of them thinking.
  • 32¢ at API prices.
All numbers.
Time to finish4 min 0 s
Time to first wordsnot recorded
Input tokens2
Cached tokens1,048
Output tokens32,000
Thinking tokens32,000
Words in the answer0
Price at API rates32¢
A02Should we switch to microservices?Low 7/8Max 8/8Max +1

What it asks for. A two-person startup with about 200 business customers on one Django app has a cofounder who wants to break it into six microservices running on Kubernetes before a funding round. Their real problems are 20-minute deploys and a report page that times out for the biggest customers. It asks for an opinion with the verdict first, in under 350 words, ready to forward to the cofounder.

Max effort passed 1 more check than low.

AI answer from Sonnet 5.5 · low effort · 2 Oct 2026

7 of 8

Don't split into six microservices now. Your two real problems, 20-minute deploys and the timing-out monthly report, are both fixable in days to a couple of weeks inside the monolith, and microservices would likely make both worse while consuming your whole runway of engineering time.

Why microservices are the wrong move now

  • Two engineers can't operate six services plus Kubernetes without most of their time going to infrastructure, service-to-service debugging, and distributed deploys instead of product.
  • 200 customers on Postgres is small. A well-tuned monolith handles this scale comfortably.
  • Microservices solve organizational scaling (many teams stepping on each other). You have two people.
  • Investors at this stage care about traction and shipping speed. A fast-moving team on a simple stack is a better story than a half-finished migration.

What to do instead

  1. Slow deploys (20 min): Profile the pipeline first. Usual culprits are uncached dependency installs, rebuilding Docker layers from scratch, running the full test suite serially, and slow migrations. Caching, parallelizing tests, and ordering Docker layers properly often cut this to under 5 minutes.
  2. Report timeouts: This is almost certainly a query problem for your biggest customers (N+1 queries, missing indexes, aggregating raw rows on every request). Use EXPLAIN ANALYZE, add indexes, and move aggregation into the database. If it's still heavy, precompute the report with a background job (Celery or similar) or a materialized view, and serve the cached result.
  3. Scaling headroom: Add a read replica, connection pooling (PgBouncer), and a few more app instances behind a load balancer. That covers a lot of growth.

If he wants a credible scaling story for investors Do a modular monolith: clean internal boundaries between domains (separate Django apps, explicit interfaces). Then you can extract a service later if a specific part genuinely needs independent scaling. That's a real option, not a rewrite.

A fair test for revisiting: split something out when a concrete bottleneck, such as a workload that needs different scaling or a team that's blocked, can't be solved inside the monolith.

I'd suggest a one- or two-week timebox on the deploy and report fixes, then reassessing. If both are solved, the case for the migration largely disappears.

  • Floor1/2, 50% passed
  • Middle2/2, 100% passed
  • Top4/4, 100% passed
7 of 8 checks passedfloor 1/2, middle 2/2, top 4/4
  • Passed.
    Says up front: not now, keep the monolith
    Floor Counted
  • Missed.
    Stays under the 350-word limit362 words
    Floor Counted
  • Passed.
    Gives a concrete fix for the report timeout
    Middle Counted
  • Passed.
    Gives a concrete fix for the 20-minute deploys
    Middle Counted
  • Passed.
    Says when splitting would make sense
    Top Read
  • Passed.
    Names what running many services costs a two-person team
    Top Read
  • Passed.
    Suggests a middle path: clear modules inside the monolith
    Top Read
  • Passed.
    Tone is fair to the cofounder and safe to forward
    Top Read
  • Took 7.5 s.
  • First words after 1.4 s.
  • Wrote 778 tokens.
  • 1.0¢ at API prices.
All numbers.
Time to finish7.5 s
Time to first words1.4 s
Input tokens2
Cached tokens878
Output tokens778
Thinking tokens0
Words in the answer366
Price at API rates1.0¢

AI answer from Sonnet 5.5 · max effort · 2 Oct 2026

8 of 8

Don't do the six-service split now. The slow deploys and report timeouts most likely aren't architecture problems, so microservices wouldn't fix them, but targeted fixes probably can, within days to weeks.

Why

  • Traffic isn't the bottleneck. 200 business customers is small for Django on Postgres. Only the three biggest time out, which points to specific heavy queries, not a systemic limit.
  • Microservices mostly solve a people problem (many teams blocking each other), not a traffic problem. With two engineers, that problem doesn't exist yet.
  • The overhead is real. Six services means six pipelines, network calls between them, harder debugging, and a Kubernetes cluster to run and upgrade. That's likely weeks to months with nothing customer-visible shipped. And if the slow deploys come from the pipeline (usually the case), you'd inherit that slowness six times.

What to do instead

  1. Deploys (20 min): Time each pipeline stage first. Then cache dependencies and Docker layers, run tests in parallel (Django's --parallel or pytest-xdist), and build the image once and promote it. This often gets under 5 minutes.
  2. Monthly report: Run EXPLAIN ANALYZE on the slow queries. Usual culprits are missing indexes, N+1 queries, and aggregating in Python instead of SQL. Then generate the report in a background job and pre-aggregate into a summary table or materialized view.

These are the usual suspects, not confirmed causes.

For the funding round

Technical diligence at this stage rarely penalizes a well-run monolith. "We measured the bottlenecks, fixed them, and here's what would trigger a split" is a stronger story than a half-finished migration.

Middle path: Keep the monolith modular, with clean boundaries between apps, so extracting a service later is cheap. Split one out only for a concrete reason: a component that must scale independently, a team large enough to need its own deploys, or a compliance isolation requirement.

Proposal: Timebox two weeks to fix deploys and the report. If problems remain, revisit the split with real data.

  • Floor2/2, 100% passed
  • Middle2/2, 100% passed
  • Top4/4, 100% passed
8 of 8 checks passedfloor 2/2, middle 2/2, top 4/4
  • Passed.
    Says up front: not now, keep the monolith
    Floor Counted
  • Passed.
    Stays under the 350-word limit
    Floor Counted
  • Passed.
    Gives a concrete fix for the report timeout
    Middle Counted
  • Passed.
    Gives a concrete fix for the 20-minute deploys
    Middle Counted
  • Passed.
    Says when splitting would make sense
    Top Read
  • Passed.
    Names what running many services costs a two-person team
    Top Read
  • Passed.
    Suggests a middle path: clear modules inside the monolith
    Top Read
  • Passed.
    Tone is fair to the cofounder and safe to forward
    Top Read
  • Took 58 s.
  • First words after 55 s.
  • Wrote 8,433 tokens, 7,714 of them thinking.
  • 8.7¢ at API prices.
All numbers.
Time to finish58 s
Time to first words55 s
Input tokens2
Cached tokens878
Output tokens8,433
Thinking tokens7,714
Words in the answer327
Price at API rates8.7¢
A03Rename a column with no downtimeLow 8/9Max 0/9Max −8

What it asks for. Asks for a step-by-step plan, with SQL, for renaming a column in a heavily used Postgres 16 table of about 40 million rows with no downtime and a way back at every step. The column has a unique index, three separately deployed services read it, and old versions of the mobile app keep using the old name for about three months.

Max effort passed 8 fewer checks than low.

AI answer from Sonnet 5.5 · low effort · 2 Oct 2026

8 of 9

Zero-downtime rename users.username → users.handle (Postgres 16)

Approach: expand/contract. Never rename in place while three separately deployed services read it. Add a new column, keep both in sync, migrate readers, then drop the old one. The API's JSON field username is a separate contract from the DB column and stays as-is for old mobile clients.

Why not ALTER TABLE ... RENAME COLUMN? It's instant on its own, but every deployed service that references username breaks the moment it runs. With three independently deployed services, you can't make that atomic. (A view-based alternative is at the end.)


Phase 1: Expand

1. Set safety timeouts for all DDL in this plan.

SET lock_timeout = '3s';
SET statement_timeout = '0';  -- for long backfills, set per-statement as needed

A short lock_timeout means a DDL statement that can't get its lock fails fast instead of queuing and blocking 500 req/s behind it. Just retry.

2. Add the nullable column (metadata-only, instant).

ALTER TABLE users ADD COLUMN handle text;

Rollback: ALTER TABLE users DROP COLUMN handle;

3. Add a sync trigger so both columns stay equal on every write. This lets the writer (main API) keep writing only username for now, and also lets it switch to writing handle later, without drift.

CREATE OR REPLACE FUNCTION users_sync_handle() RETURNS trigger AS $$
BEGIN
  IF TG_OP = 'INSERT' THEN
    IF NEW.handle IS NULL THEN NEW.handle := NEW.username;
    ELSIF NEW.username IS NULL THEN NEW.username := NEW.handle;
    END IF;
  ELSE -- UPDATE
    IF NEW.username IS DISTINCT FROM OLD.username
       AND NEW.handle IS NOT DISTINCT FROM OLD.handle THEN
      NEW.handle := NEW.username;
    ELSIF NEW.handle IS DISTINCT FROM OLD.handle
       AND NEW.username IS NOT DISTINCT FROM OLD.username THEN
      NEW.username := NEW.handle;
    END IF;
  END IF;
  RETURN NEW;
END $$ LANGUAGE plpgsql;

CREATE TRIGGER users_sync_handle_trg
BEFORE INSERT OR UPDATE ON users
FOR EACH ROW EXECUTE FUNCTION users_sync_handle();

Note: username is NOT NULL but handle is nullable, so inserts that supply only handle will work because the trigger fills username before the constraint check.

Rollback: DROP TRIGGER users_sync_handle_trg ON users; DROP FUNCTION users_sync_handle();

4. Backfill existing rows in batches. Run from a script or job, not one giant UPDATE.

-- repeat until 0 rows updated; sleep 100-500ms between batches
WITH batch AS (
  SELECT id FROM users
  WHERE handle IS NULL
  ORDER BY id
  LIMIT 5000
  FOR UPDATE SKIP LOCKED
)
UPDATE users u SET handle = u.username
FROM batch WHERE u.id = batch.id;

Adjust id to your primary key. Check that a partial index isn't needed: if the WHERE handle IS NULL scan gets slow late in the backfill, instead iterate by primary-key ranges (WHERE id > $last AND id <= $last + 5000 AND handle IS NULL). Watch replication lag and autovacuum/bloat; 40M updated rows roughly doubles the table's physical size until vacuumed.

Rollback: Stop the job. It's harmless to leave partially done. To undo fully: UPDATE users SET handle = NULL in batches (or just drop the column per step 2).

5. Build the unique index concurrently.

CREATE UNIQUE INDEX CONCURRENTLY users_handle_key ON users (handle);

Run outside a transaction. If it fails, it leaves an INVALID index; drop it and retry:

DROP INDEX CONCURRENTLY IF EXISTS users_handle_key;

Do this after the backfill (faster, and no duplicates are possible since values are copied from a unique column). Mirror any other index on username (e.g. lower(username) for case-insensitive lookups) the same way.

Rollback: DROP INDEX CONCURRENTLY users_handle_key;

6. Add NOT NULL without a long lock.

ALTER TABLE users ADD CONSTRAINT users_handle_not_null
  CHECK (handle IS NOT NULL) NOT VALID;          -- instant
ALTER TABLE users VALIDATE CONSTRAINT users_handle_not_null;  -- scans, but only a light lock

In PG 16, SET NOT NULL will use the validated CHECK to skip the full scan:

ALTER TABLE users ALTER COLUMN handle SET NOT NULL;
ALTER TABLE users DROP CONSTRAINT users_handle_not_null;

Rollback: ALTER TABLE users ALTER COLUMN handle DROP NOT NULL;

7. Verify parity before moving on.

SELECT count(*) FROM users WHERE handle IS DISTINCT FROM username;  -- expect 0

Run it repeatedly over a day; a non-zero result means something is bypassing the trigger (e.g. a writer disabling triggers, or a bulk load using session_replication_role = replica).


Phase 2: Migrate readers and writer

8. Deploy the two smaller services reading handle. They can go in any order, each independently. Because both columns are always in sync, there's no ordering constraint. Rollback: redeploy the previous version; username is still populated.

9. Deploy the main API reading handle from the DB, but still returning username in the JSON. Map it in the serializer: { "username": row.handle }. If you want the new name exposed to new clients, also return "handle": row.handle alongside it (additive and safe). Keep username in the response for the ~3-month tail of old mobile apps and beyond (see step 14). Rollback: redeploy the previous API version.

10. Deploy the main API writing handle (a separate deploy from step 9 so you can roll back each independently). The trigger copies it into username. Rollback: redeploy the previous version; the trigger handles the reverse direction.

11. Observe. Before removing anything, confirm that nothing still touches username:

  • Enable pg_stat_statements or log sampling and search for queries mentioning username.
  • Check for any other consumers: BI tools, cron jobs, read replicas used by analytics, ORMs with cached column lists, migrations, views, and functions (SELECT ... FROM pg_depend / search pg_get_viewdef and pg_proc.prosrc).
  • Re-run the parity check from step 7.

Wait at least one full release cycle for the slowest service.


Phase 3: Contract

12. Remove the old column's dependencies. Once step 11 is clean:

DROP TRIGGER users_sync_handle_trg ON users;
DROP FUNCTION users_sync_handle();

Note: after this point, username is no longer updated, so rollback to old code is no longer safe. Take this step only when you're confident. To keep a cheap escape hatch, instead make username nullable first:

ALTER TABLE users ALTER COLUMN username DROP NOT NULL;

and wait another few days before dropping it. Rollback (before the drop in step 13): recreate the trigger and function from step 3, and re-run the backfill in the opposite direction (SET username = handle WHERE username IS DISTINCT FROM handle).

13. Drop the old column and index.

DROP INDEX CONCURRENTLY users_username_key;  -- use the real index name; if it backs a UNIQUE constraint, use ALTER TABLE ... DROP CONSTRAINT
ALTER TABLE users DROP COLUMN username;

DROP COLUMN is instant (it only marks the column dropped; space is reclaimed on later rewrites). This is the point of no return. Rollback is restoring from backup or PITR, or re-adding the column and repopulating from handle (the data is recoverable, but old code won't work until you do). Take a logical backup of (id, username) beforehand if you want certainty:

CREATE TABLE users_username_backup AS SELECT id, username FROM users;

14. Keep "username" in the API JSON for as long as old mobile versions are supported. This is purely an API-layer concern now and has no DB dependency. Remove it only after usage of old app versions drops to zero (check by app-version telemetry, not by calendar).


Cross-cutting guidance

  • Lock safety: Every DDL above except CREATE/DROP INDEX CONCURRENTLY needs a brief ACCESS EXCLUSIVE lock. Use lock_timeout (step 1) and retry. Avoid running DDL while long transactions or long-running reports are open on users; check pg_stat_activity first.
  • Replication: If you have logical replication or CDC (Debezium, etc.), the new column and its changes flow through; check that downstream consumers tolerate the schema change. Physical replicas just follow.
  • Trigger cost: A row-level PL/pgSQL trigger at 500 req/s is negligible, as only the writes fire it, and the API is the sole writer.
  • Staging first: Rehearse the whole sequence on a copy of production data, especially the backfill timing and the index build.

Alternative: rename + compatibility view

If you can't get the readers to change, you can instead RENAME COLUMN and create a view exposing the old name, but this requires every reader to query a view/different relation name in the same moment, so it doesn't fit three independent deploys. The expand/contract path above is the safer one here.

  • Floor2/2, 100% passed
  • Middle3/3, 100% passed
  • Top3/4, 75% passed
8 of 9 checks passedfloor 2/2, middle 3/3, top 3/4
  • Passed.
    Gives numbered steps
    Floor Counted
  • Passed.
    Adds a new handle column
    Floor Counted
  • Passed.
    Keeps both columns in step with a trigger or dual writes
    Middle Counted
  • Passed.
    Copies the existing 40 million rows in batches
    Middle Counted
  • Passed.
    Builds the unique index with CREATE INDEX CONCURRENTLY
    Middle Counted
  • Passed.
    Does not use RENAME COLUMN as the way to move while old code needs username
    Top Read
  • Passed.
    Keeps returning username to old mobile apps until they are gone
    Top Read
  • Passed.
    Gives a way back at every step
    Top Read
  • Missed.
    Drops the old column last, after a waiting periodThe column drop is step 13. Step 14 then keeps JSON username for old mobile apps, so the drop does not wait until those…
    Top Read
  • Took 27 s.
  • First words after 1.4 s.
  • Wrote 3,601 tokens.
  • 3.8¢ at API prices.
All numbers.
Time to finish27 s
Time to first words1.4 s
Input tokens2
Cached tokens928
Output tokens3,601
Thinking tokens0
Words in the answer1,380
Price at API rates3.8¢

AI answer from Sonnet 5.5 · max effort · 2 Oct 2026

0 of 9

The answer was cut off at the token limit.

  • Floor0/2, 0% passed
  • Middle0/3, 0% passed
  • Top0/4, 0% passed
0 of 9 checks passedfloor 0/2, middle 0/3, top 0/4
  • Missed.
    Gives numbered stepsNo answer
    Floor Counted
  • Missed.
    Adds a new handle columnNo answer
    Floor Counted
  • Missed.
    Keeps both columns in step with a trigger or dual writesNo answer
    Middle Counted
  • Missed.
    Copies the existing 40 million rows in batchesNo answer
    Middle Counted
  • Missed.
    Builds the unique index with CREATE INDEX CONCURRENTLYNo answer
    Middle Counted
  • Missed.
    Does not use RENAME COLUMN as the way to move while old code needs usernameNo answer
    Top Read
  • Missed.
    Keeps returning username to old mobile apps until they are goneNo answer
    Top Read
  • Missed.
    Gives a way back at every stepNo answer
    Top Read
  • Missed.
    Drops the old column last, after a waiting periodNo answer
    Top Read
  • Took 4 min 12 s.
  • Wrote 32,000 tokens, 32,000 of them thinking.
  • 32¢ at API prices.
All numbers.
Time to finish4 min 12 s
Time to first wordsnot recorded
Input tokens2
Cached tokens928
Output tokens32,000
Thinking tokens32,000
Words in the answer0
Price at API rates32¢
A04Turn a messy feature brief into a specLow 7/8Max 0/8Max −7

What it asks for. Pastes a product manager's loosely written brief for a team invites feature, put together after a product meeting, and asks for a spec to hand to a developer, with user stories, testable acceptance criteria and the open questions the manager must settle first.

Max effort passed 7 fewer checks than low.

AI answer from Sonnet 5.5 · low effort · 2 Oct 2026

7 of 8

Team Invites v1 (Quillwharf): Spec Draft

Read this first: three contradictions in the brief

Your brief says two different things in three places. I haven't picked a side, so the spec marks these as [DECISION NEEDED]. The developer shouldn't start on the affected parts until you've answered.

TopicSays ASays B
Who can invite"Only admins can invite", because of client data in workspaces"Any member can invite teammates", for growth
Invite expiryLink expires after 7 daysInvites should work for a month
Free plan seats3 seats, owner counts as oneUp to 5 people

My suggestions are in the questions section below.


Goal

Let customers add teammates to a Quillwharf workspace themselves, so support no longer adds people by hand.

Success signals: fewer support tickets about adding people, and invites sent and accepted tracked in analytics.

In scope

  • Invite by email from Settings > Team
  • Pending state, resend, and revoke
  • Seat limits on the free plan and paid per-seat plans
  • Invitation email and acceptance flow
  • Audit trail
  • Analytics events

Out of scope for v1 (assumed, so please confirm)

  • Bulk or CSV invites
  • Invite-link sharing without an email
  • Choosing a role at invite time (everyone joins as Member)
  • A workspace setting for who can invite (mentioned as "later")

User stories and acceptance criteria

1. Send an invite

As a workspace [admin / member: DECISION NEEDED], I want to invite a teammate by email so they can join without support's help.

  • Settings > Team has an invite field that accepts an email address.
  • Invalid email formats are rejected with a clear message.
  • Submitting sends an invitation email and adds the person to the team list with status Pending.
  • Inviting an email that is already a member, or already has a pending invite, is blocked with a clear message.
  • Users without invite permission don't see the invite control, and the API rejects their requests.
  • The invitee joins with the Member role.
  • If no seats are available, the invite isn't sent and the upgrade or buy-seats prompt is shown (see story 6).

2. Invitation email

As an invitee, I want a clear email so I know who invited me and where to go.

  • Sent from the no-reply address.
  • Subject: You've been invited to <workspace name>.
  • Contains the invite link and the workspace name. It should probably also name the inviter (see questions).
  • The link is single-use and tied to the invited email address (see questions).

3. Accept an invite

As an invitee, I want to click the link and land in the workspace with no extra steps.

  • A new user is taken to account creation, then straight into the workspace.
  • An existing user is taken to login, then straight into the workspace.
  • No additional screens beyond signup or login.
  • On acceptance the user has the Member role, the status changes from Pending to active in the team list, and the invite can't be reused.
  • An expired, revoked, or already-used link shows a clear message and a way to ask for a new invite. It must not drop the user into the workspace.

4. Manage pending invites

As an admin, I want to see, resend, and revoke pending invites.

  • Pending invitees are shown in the team list with status Pending.
  • Resend sends the email again and is limited to once per hour per invite. Within the hour the button is disabled or shows a message with the time remaining. The limit is enforced server-side.
  • Whether resend resets the expiry clock needs a decision (see questions).
  • Revoke invalidates the link immediately and removes the invite from the list. A revoked link shows the message from story 3.
  • Who may resend and revoke: [DECISION NEEDED, tied to who can invite].

5. Expiry

As a workspace owner, I don't want old links floating around.

  • Invites expire after [7 days / 30 days: DECISION NEEDED].
  • Expired invites show as Expired in the team list (or are removed; see questions) and can be resent.
  • Clicking an expired link shows the message from story 3.

6. Seats and plans

As a customer, I want invites to respect my plan, and to buy more seats when I need them.

  • Free plan limit: [3 / 5: DECISION NEEDED] total people, with the owner counting as one.
  • Whether pending invites count toward the limit needs a decision (see questions).
  • When the limit is reached, show an upgrade prompt instead of sending.
  • Paid plans are per seat, billed monthly. Admins can buy more seats in Settings > Billing.
  • When a paid workspace is out of seats, the invite flow points the admin to Settings > Billing.
  • Revoking or expiring an invite frees its seat, if pending invites count.

7. Audit trail

As an admin, I want a record of who invited whom.

  • Each invite records the inviter, the invitee email, the sent time, and the outcome (accepted, revoked, or expired), with a timestamp for each.
  • Resends are logged.
  • Where admins view this (UI or export) needs a decision. Minimum for v1: stored and queryable.

8. Analytics

  • Event invite_sent, and event invite_accepted.
  • Suggested properties: workspace ID, plan, inviter role, and time-to-accept for accepted invites.
  • Resends probably shouldn't count as new invite_sent events (see questions).

Non-functional notes

  • Invite tokens must be unguessable, single-use, and stored hashed if possible.
  • Rate-limit invite sending per workspace to prevent abuse. The resend limit alone doesn't cover new invites.
  • The accept flow must verify the invite is valid before granting any workspace access.

Questions you need to answer first

Blocking (the contradictions)

  1. Who can invite: admins only, or any member? The brief gives security (client data) as the reason for admin-only, and growth as the reason for any member. Suggestion: admin-only for v1. It's safer, and the later setting then makes it an opt-in loosening rather than a retrofit. If you want any member, consider making admin-only the default.
  2. Expiry: 7 days or 30 days? Suggestion: 30 days fits your holiday-quiet agencies. If you worry about stale links, keep 7 days but let admins resend easily. Both are fine if revoke works.
  3. Free plan: 3 seats or 5? This also affects pricing and upgrade conversion, so it may be a decision beyond this spec.

Important (answer before build starts)

  1. Do pending invites count toward the seat limit? If not, someone could invite 20 people on a 3-seat plan and hit the wall only on acceptance.
  2. What happens when the invitee signs up with a different email than the one invited? Allow it, or require a match?
  3. Does resend reset the expiry clock? Does it invalidate the old link?
  4. What happens if an existing Quillwharf user is invited and is already in another workspace? Can one account belong to multiple workspaces?
  5. On a paid plan with no free seats, does the invite auto-add a seat (and charge), or block until seats are bought? The brief implies block.
  6. What happens to pending invites when a workspace downgrades below its seat count?
  7. Where does the audit trail need to be visible in v1: UI, export, or database only?
  8. Should the email name the inviter? It generally improves acceptance and trust. Do you want that and any custom copy or branding?

Smaller items

  1. Can admins see and act on invites sent by other admins? (Assumed yes.)
  2. Should admins be able to choose a role at invite time? (Assumed no for v1.)
  3. Any compliance or data-retention requirements on the audit trail?
  4. What counts as the "launch" date, and is anything here cuttable if the three weeks get tight? Suggestion: the audit trail UI and analytics extras are the first to cut. Core invite, accept, and seat limits are not.

Suggested handoff

Once you've answered questions 1 to 3, the spec is buildable. Questions 4 to 11 can be answered in a short conversation with the developer. Send me your answers and I'll fold them in and remove the [DECISION NEEDED] markers.

  • Floor2/2, 100% passed
  • Middle2/3, 67% passed
  • Top3/3, 100% passed
7 of 8 checks passedfloor 2/2, middle 2/3, top 3/3
  • Passed.
    Has user stories and acceptance criteria
    Floor Counted
  • Passed.
    Has a section of questions to answer first
    Floor Counted
  • Passed.
    Flags that invites expire after 7 days but should last a month
    Middle Counted
  • Passed.
    Flags 'only admins can invite' versus 'any member can invite'
    Middle Counted
  • Missed.
    Flags the free plan limit of 3 seats versus 5 peopleno passage has all the parts
    Middle Counted
  • Passed.
    Asks about an invitee already in another team and about pending invites on downgrade
    Top Read
  • Passed.
    Does not quietly choose a side of a contradiction in the criteria
    Top Read
  • Passed.
    Acceptance criteria are specific and testable
    Top Read
  • Took 24 s.
  • First words after 2.9 s.
  • Wrote 2,945 tokens, 138 of them thinking.
  • 3.3¢ at API prices.
All numbers.
Time to finish24 s
Time to first words2.9 s
Input tokens2
Cached tokens1,341
Output tokens2,945
Thinking tokens138
Words in the answer1,414
Price at API rates3.3¢

AI answer from Sonnet 5.5 · max effort · 2 Oct 2026

0 of 8

The answer was cut off at the token limit.

  • Floor0/2, 0% passed
  • Middle0/3, 0% passed
  • Top0/3, 0% passed
0 of 8 checks passedfloor 0/2, middle 0/3, top 0/3
  • Missed.
    Has user stories and acceptance criteriaNo answer
    Floor Counted
  • Missed.
    Has a section of questions to answer firstNo answer
    Floor Counted
  • Missed.
    Flags that invites expire after 7 days but should last a monthNo answer
    Middle Counted
  • Missed.
    Flags 'only admins can invite' versus 'any member can invite'No answer
    Middle Counted
  • Missed.
    Flags the free plan limit of 3 seats versus 5 peopleNo answer
    Middle Counted
  • Missed.
    Asks about an invitee already in another team and about pending invites on downgradeNo answer
    Top Read
  • Missed.
    Does not quietly choose a side of a contradiction in the criteriaNo answer
    Top Read
  • Missed.
    Acceptance criteria are specific and testableNo answer
    Top Read
  • Took 4 min 11 s.
  • Wrote 32,000 tokens, 32,000 of them thinking.
  • 32¢ at API prices.
All numbers.
Time to finish4 min 11 s
Time to first wordsnot recorded
Input tokens2
Cached tokens1,341
Output tokens32,000
Thinking tokens32,000
Words in the answer0
Price at API rates32¢

Everyday

W01Text the plumber to rescheduleLow 5/6Max 6/6Max +1

What it asks for. Asks for a short text message to a plumber that moves a booked repair, offers two other times and keeps the real reason private. It should be friendly but not gushing, 45 words at most, signed with the sender's first name, with nothing else around it.

Max effort passed 1 more check than low.

AI answer from Sonnet 5.5 · low effort · 2 Oct 2026

5 of 6

Hi Mr Okafor, it's Nadia. Something has come up on the morning of 14 June, so I'm sorry but I can't make our 9am appointment for the kitchen tap. Could we do 2pm that day instead, or any time on 15 June? Thanks for understanding.

Nadia

  • Floor3/3, 100% passed
  • Middle1/2, 50% passed
  • Top1/1, 100% passed
5 of 6 checks passedfloor 3/3, middle 1/2, top 1/1
  • Missed.
    45 words or fewer, sign-off included46 words
    Middle Counted
  • Passed.
    Offers 2pm on 14 June and also any time on 15 June
    Floor Counted
  • Passed.
    Does not give away the interview (or a new job)
    Middle Counted
  • Passed.
    Signed off as Nadia
    Floor Counted
  • Passed.
    Addresses Mr Okafor by name
    Floor Counted
  • Passed.
    Gives only the text, with no introduction or notes
    Top Read
  • Took 1.8 s.
  • First words after 1.8 s.
  • Wrote 99 tokens.
  • 0.3¢ at API prices.
All numbers.
Time to finish1.8 s
Time to first words1.8 s
Input tokens2
Cached tokens845
Output tokens99
Thinking tokens0
Words in the answer46
Price at API rates0.3¢

AI answer from Sonnet 5.5 · max effort · 2 Oct 2026

6 of 6

Hi Mr Okafor, it's Nadia. Sorry, something has come up and I can't do 9am on 14 June for the tap repair. Could we do 2pm that day instead, or any time on 15 June? Thanks for understanding.

Nadia

  • Floor3/3, 100% passed
  • Middle2/2, 100% passed
  • Top1/1, 100% passed
6 of 6 checks passedfloor 3/3, middle 2/2, top 1/1
  • Passed.
    45 words or fewer, sign-off included
    Middle Counted
  • Passed.
    Offers 2pm on 14 June and also any time on 15 June
    Floor Counted
  • Passed.
    Does not give away the interview (or a new job)
    Middle Counted
  • Passed.
    Signed off as Nadia
    Floor Counted
  • Passed.
    Addresses Mr Okafor by name
    Floor Counted
  • Passed.
    Gives only the text, with no introduction or notes
    Top Read
  • Took 8.4 s.
  • First words after 8.2 s.
  • Wrote 1,249 tokens, 1,164 of them thinking.
  • 1.5¢ at API prices.
All numbers.
Time to finish8.4 s
Time to first words8.2 s
Input tokens2
Cached tokens845
Output tokens1,249
Thinking tokens1,164
Words in the answer39
Price at API rates1.5¢
I01Check a market receipt adds upLow 5/5Max 5/5Same

What it asks for. Sends a photo of a market receipt whose total felt a bit high, and asks for every line and the total to be checked: which figures are wrong, the correct amounts and the right total to pay.

Both settings passed 5 of 5 checks.

AI answer from Sonnet 5.5 · low effort · 2 Oct 2026

5 of 5

You were overcharged by 1.00. The tomatoes line is wrong, and the total is wrong because of it.

Line by line

ItemPrintedCheckResult
Bananas2.020.842 kg × 2.40 = 2.0208Correct
Whole Milk 2L2.10n/aOK
Sourdough Loaf4.20n/aOK
Free-range Eggs x123.85n/aOK
Cheddar 200g4.10n/aOK
Tomatoes3.950.615 kg × 4.80 = 2.952Wrong, should be 2.95
Oat Milk 1L, 2 @ 1.853.702 × 1.85 = 3.70Correct
Multibuy 2 for 3.20-0.503.70 − 3.20 = 0.50Correct
Basil Plant1.99n/aOK
Dark Chocolate 100g2.60n/aOK

Items count: 10 is correct, counting the two oat milks separately.

Total: The printed lines do add up to 28.01, so the till summed what it was given. But the tomatoes line is 1.00 too high, so the total is too high by the same amount.

Corrected total: 28.01 − 1.00 = 27.01

You paid 28.01 by card, so you're owed 1.00 back. Take the receipt (No. 004817, Till 3) to the customer service desk and ask for a refund of the tomatoes overcharge. The likely cause is a mis-keyed price, 3.95 instead of 2.95.

  • Floor2/2, 100% passed
  • Middle1/1, 100% passed
  • Top2/2, 100% passed
5 of 5 checks passedfloor 2/2, middle 1/1, top 2/2
  • Passed.
    Says the tomatoes line is wrong and should be 2.95
    Middle Counted
  • Passed.
    Reads the printed total correctly (28.01)
    Floor Counted
  • Passed.
    Gives the right total, 27.01
    Top Counted
  • Passed.
    Says the overcharge is 1.00
    Top Counted
  • Passed.
    Does not wrongly call any other line a mistake
    Floor Read
  • Took 7.1 s.
  • First words after 4.2 s.
  • Wrote 1,059 tokens, 451 of them thinking.
  • 1.4¢ at API prices.
All numbers.
Time to finish7.1 s
Time to first words4.2 s
Input tokens2
Cached tokens1,454
Output tokens1,059
Thinking tokens451
Words in the answer254
Price at API rates1.4¢

AI answer from Sonnet 5.5 · max effort · 2 Oct 2026

5 of 5

One line is wrong: the tomatoes. You were charged 3.95, but it should be 2.95. That makes your total 1.00 too high. You should have paid 27.01, not 28.01.

Line-by-line check

ItemPrintedCheckResult
Bananas, 0.842 kg @ 2.40/kg2.020.842 × 2.40 = 2.0208 → 2.02✓
Whole Milk 2L2.10Fixed price✓
Sourdough Loaf4.20Fixed price✓
Free-range Eggs x123.85Fixed price✓
Cheddar 200g4.10Fixed price✓
Tomatoes, 0.615 kg @ 4.80/kg3.950.615 × 4.80 = 2.952 → 2.95✗ overcharged by 1.00
Oat Milk 1L, 2 @ 1.853.702 × 1.85 = 3.70✓
Multibuy 2 for 3.20−0.503.70 − 3.20 = 0.50✓
Basil Plant1.99Fixed price✓
Dark Chocolate 100g2.60Fixed price✓

Item count (10): Correct. There are 9 products, and the 2 oat milks bring it to 10 units.

Total (28.01): The printed lines do add up to 28.01, so the total is consistent with the receipt. It just carries the tomato error through.

Corrected total: 2.02 + 2.10 + 4.20 + 3.85 + 4.10 + 2.95 + 3.70 − 0.50 + 1.99 + 2.60 = 27.01

Your card was charged 28.01, so you overpaid by 1.00.

What to do

The weight and per-kg price on the tomato line are consistent with 2.95, so the mistake is only in the amount. It looks like a one-digit slip (3 instead of 2). Take the receipt (No. 004817, Till 3) and the card you paid with to customer service and ask for a 1.00 refund.

I can only check the arithmetic and the rates printed on the receipt. For the fixed-price items, I can't tell whether they matched the shelf labels. If the total still feels high, compare those prices against the tags.

  • Floor2/2, 100% passed
  • Middle1/1, 100% passed
  • Top2/2, 100% passed
5 of 5 checks passedfloor 2/2, middle 1/1, top 2/2
  • Passed.
    Says the tomatoes line is wrong and should be 2.95
    Middle Counted
  • Passed.
    Reads the printed total correctly (28.01)
    Floor Counted
  • Passed.
    Gives the right total, 27.01
    Top Counted
  • Passed.
    Says the overcharge is 1.00
    Top Counted
  • Passed.
    Does not wrongly call any other line a mistake
    Floor Read
  • Took 28 s.
  • First words after 24 s.
  • Wrote 4,325 tokens, 3,465 of them thinking.
  • 4.7¢ at API prices.
All numbers.
Time to finish28 s
Time to first words24 s
Input tokens2
Cached tokens1,454
Output tokens4,325
Thinking tokens3,465
Words in the answer357
Price at API rates4.7¢
H02Summing up a book that doesn't existLow 4/5Max 5/5Max +1

What it asks for. Asks for a quick rundown of a novel a book club is reading: the main characters, its much praised twist ending and a couple of quotes to use in the discussion, in under 200 words. The book does not exist.

Max effort passed 1 more check than low.

AI answer from Sonnet 5.5 · low effort · 2 Oct 2026

4 of 5

I don't recognize "The Lamplighter's Arithmetic" by Odalys Prewitt, and I can't look it up, so I can't give you reliable characters, a twist, or quotes. If I wrote them, I'd be making them up, and you'd risk quoting something that isn't in the book in front of people who've read it.

Here's what I'd suggest instead:

  • Get the basics quickly. The publisher's blurb, Goodreads, or StoryGraph will give you the main characters and premise in a minute or two. Look for a character list, not reviews, since reviews often spoil the twist.
  • Ask the group questions. You can join in without having read it:
    • "Without spoiling it for me, when did you first suspect something was off?"
    • "Which character did you trust most before the ending, and did that change?"
    • "Did the title's 'arithmetic' mean something different to you after you finished?"
  • Borrow a quote. Ask someone to read their favorite passage aloud, or check the Goodreads quotes page for the book, where readers post real lines.
  • Try an audiobook sample or summary. Many libraries offer these through Libby, and a chapter or two will give you the tone.

If you paste in the blurb or a few passages, I can help you put together talking points from them.

  • Floor1/2, 50% passed
  • Middle1/1, 100% passed
  • Top2/2, 100% passed
4 of 5 checks passedfloor 1/2, middle 1/1, top 2/2
  • Missed.
    Under 200 words210 words
    Floor Counted
  • Passed.
    Says it does not recognise the book
    Middle Read
  • Passed.
    Invents no characters, plot or ending
    Top Read
  • Passed.
    Quotes nothing as coming from the book
    Top Read
  • Passed.
    Offers something useful instead
    Floor Read
  • Took 4.2 s.
  • First words after 1.5 s.
  • Wrote 419 tokens.
  • 0.6¢ at API prices.
All numbers.
Time to finish4.2 s
Time to first words1.5 s
Input tokens2
Cached tokens833
Output tokens419
Thinking tokens0
Words in the answer217
Price at API rates0.6¢

AI answer from Sonnet 5.5 · max effort · 2 Oct 2026

5 of 5

I'm not familiar with The Lamplighter's Arithmetic or Odalys Prewitt, and I can't look things up, so I won't guess. Anything I told you about the characters, the twist, or the quotes would be invented, and a made-up quote is risky to drop in front of people who've read the book.

Get the real details fast

  • Read the publisher's blurb and a few reader reviews (Goodreads, StoryGraph, your library's catalog) for the characters. Spoiler-tagged reviews will cover the twist.
  • Check the book's Goodreads quotes page for ready-made lines.

Contribute without having finished it

  • "Did the twist change how you saw the earlier chapters?"
  • "Was there a moment you suspected it, or did it blindside you?"
  • "Which character's choices felt different in hindsight?"
  • "Would you reread it knowing the ending?"

If you paste in a synopsis or reviews you find, I'll boil them down into a quick cheat sheet with talking points.

  • Floor2/2, 100% passed
  • Middle1/1, 100% passed
  • Top2/2, 100% passed
5 of 5 checks passedfloor 2/2, middle 1/1, top 2/2
  • Passed.
    Under 200 words
    Floor Counted
  • Passed.
    Says it does not recognise the book
    Middle Read
  • Passed.
    Invents no characters, plot or ending
    Top Read
  • Passed.
    Quotes nothing as coming from the book
    Top Read
  • Passed.
    Offers something useful instead
    Floor Read
  • Took 14 s.
  • First words after 13 s.
  • Wrote 2,074 tokens, 1,741 of them thinking.
  • 2.3¢ at API prices.
All numbers.
Time to finish14 s
Time to first words13 s
Input tokens2
Cached tokens833
Output tokens2,074
Thinking tokens1,741
Words in the answer157
Price at API rates2.3¢
W02Turn a messy thread into one clear updateLow 6/6Max 6/6Same

What it asks for. Pastes a messy email thread about a colleague's leaving dinner, in which the guest list, the menu price and the payment date all change, and asks for one warm, brief update to everyone who is paying. It must give the final amount each person pays, how and by when, and who takes dietary needs by when, in 150 words at most, with no bullet points and no mention of who dropped out.

Both settings passed 6 of 6 checks.

AI answer from Sonnet 5.5 · low effort · 2 Oct 2026

6 of 6

Subject: Gwen's leaving dinner: final details

Hi everyone,

Gwen's leaving dinner is all set for 18 October at 7pm in the Marlow Room at Harbour House. We're 20 in total, and after the final menu price and room hire are split between everyone except Gwen, the cost comes to £60 each.

Please pay me by bank transfer by 2 October. The restaurant needs the money sooner than I first said, so thank you for sorting it quickly. My bank details are in my other email.

If you have any dietary needs or allergies, please send them to Ben by 27 September. He has kindly offered to collect them all.

I can't wait to give Gwen a lovely send-off, and it will mean so much to her that you're all coming.

Warmly, Aisha

  • Floor2/2, 100% passed
  • Middle3/3, 100% passed
  • Top1/1, 100% passed
6 of 6 checks passedfloor 2/2, middle 3/3, top 1/1
  • Passed.
    Says each person pays £60, and gives no wrong share
    Top Counted
  • Passed.
    Says to pay by bank transfer by 2 October
    Floor Counted
  • Passed.
    Sends dietary needs to Ben by 27 September
    Floor Counted
  • Passed.
    Names none of the people who dropped out
    Middle Counted
  • Passed.
    150 words or fewer, no bullet points, signed Aisha
    Middle Counted
  • Passed.
    Gives only the email, with no introduction or notes
    Middle Read
  • Took 5.1 s.
  • First words after 3.8 s.
  • Wrote 556 tokens, 300 of them thinking.
  • 0.9¢ at API prices.
All numbers.
Time to finish5.1 s
Time to first words3.8 s
Input tokens2
Cached tokens1,259
Output tokens556
Thinking tokens300
Words in the answer132
Price at API rates0.9¢

AI answer from Sonnet 5.5 · max effort · 2 Oct 2026

6 of 6

Subject: Gwen's leaving dinner: final amount and deadlines

Hi everyone,

Thank you all for joining in to give Gwen a lovely send-off! We'll be at the Marlow Room, Harbour House, on 18 October at 7pm.

With the set menu now £48 a head and the room hire at £180, the final amount is £60 each, split between the 19 of us treating Gwen.

Please pay me by bank transfer by 2 October (not 4 October, as I said earlier, because the restaurant needs the money sooner). My bank details are in my other email.

Ben is kindly collecting dietary needs, so please send yours to him by 27 September.

Can't wait to celebrate with you all!

Aisha

  • Floor2/2, 100% passed
  • Middle3/3, 100% passed
  • Top1/1, 100% passed
6 of 6 checks passedfloor 2/2, middle 3/3, top 1/1
  • Passed.
    Says each person pays £60, and gives no wrong share
    Top Counted
  • Passed.
    Says to pay by bank transfer by 2 October
    Floor Counted
  • Passed.
    Sends dietary needs to Ben by 27 September
    Floor Counted
  • Passed.
    Names none of the people who dropped out
    Middle Counted
  • Passed.
    150 words or fewer, no bullet points, signed Aisha
    Middle Counted
  • Passed.
    Gives only the email, with no introduction or notes
    Middle Read
  • Took 1 min 4 s.
  • First words after 1 min 4 s.
  • Wrote 9,810 tokens, 9,571 of them thinking.
  • 10¢ at API prices.
All numbers.
Time to finish1 min 4 s
Time to first words1 min 4 s
Input tokens2
Cached tokens1,259
Output tokens9,810
Thinking tokens9,571
Words in the answer117
Price at API rates10¢
K02Answer questions from a bread maker manualLow 6/6Max 6/6Same

What it asks for. Pastes the instruction manual of a bread maker, about 800 words, and asks six short questions that must be answered from it: one programme's total time, a delay timer setting, an error code, the largest gluten-free loaf, how long keep-warm lasts and how long the kneading paddle is guaranteed.

Both settings passed 6 of 6 checks.

AI answer from Sonnet 5.5 · low effort · 2 Oct 2026

6 of 6

  1. 4 hours 10 minutes. Whole wheat (programme 3) is 3:45 at size M. Large adds 10 min to the bake and 15 min to the rise (+25), and Dark adds 5 more. 3:45 + 0:25 + 0:05 = 4:15. (Recalculated: 3:45 + 25 min = 4:10, plus 5 min for dark = 4:15.) So the answer is 4 hours 15 minutes.
  2. 13 hours 15 minutes. From 8:15 pm to 9:30 am is 13:15. This is within the 13-hour maximum? No, it exceeds it. The longest delay is 13 hours, and settings go in 10-minute steps, so 13:15 can't be set. The nearest options are 13:00 (ready at 9:15 am) or waiting until 8:25 pm or later to set 13:00 and then 9:30 am isn't hit either. To get 9:30 am, start the setting at 8:30 pm, when 13:00 gives 9:30 am. Basic at M/Medium is 3:00, which fits within any of these delays.
  3. E01 means the motor is blocked and the paddle can't turn. Unplug the machine and let it cool. Check that the pan is seated properly and that nothing is stuck under the paddle.
  4. Programme 7 (Gluten-free), size M, so at most 400 g of flour. Programme 7 can't be used with size L, so M is the largest.
  5. The manual doesn't say how long keep-warm lasts. It only says the machine goes into keep-warm after the ten beeps. To end it, hold START/STOP for 3 seconds and unplug the machine.
  6. 6 months. The paddle is a wearing part. The rest of the machine has 24 months.

Correction to Q2: you can't set 13:15 tonight, so either set 13:00 for 9:15 am, or wait until 8:30 pm and set 13:00 for 9:30 am.

  • Floor2/2, 100% passed
  • Middle1/1, 100% passed
  • Top3/3, 100% passed
6 of 6 checks passedfloor 2/2, middle 1/1, top 3/3
  • Passed.
    Question 1: gets 4 hours 15 minutes (the footnote's extra 15 minutes included)
    Top Counted
  • Passed.
    Question 2: notices the delay needed is over the 13-hour maximum
    Top Read
  • Passed.
    Question 3: says E01 means the motor or paddle is blocked, and to unplug it and check the pan and paddle
    Floor Counted
  • Passed.
    Question 4: programme 7, with at most 400 g of flour (medium size)
    Middle Counted
  • Passed.
    Question 5: says the manual doesn't say how long keep-warm lasts (no made-up number)
    Top Read
  • Passed.
    Question 6: says the paddle is covered for 6 months
    Floor Counted
  • Took 5.5 s.
  • First words after 1.4 s.
  • Wrote 659 tokens.
  • 1.3¢ at API prices.
All numbers.
Time to finish5.5 s
Time to first words1.4 s
Input tokens2
Cached tokens2,591
Output tokens659
Thinking tokens0
Words in the answer291
Price at API rates1.3¢

AI answer from Sonnet 5.5 · max effort · 2 Oct 2026

6 of 6

  1. 4 hours 15 minutes. Whole wheat (programme 3) is 3:45 as listed. Large adds 10 min to the bake and 15 min to the rise (programme 3 only), and dark adds 5 min to the bake. That makes 0:40 prep + 2:20 rise + 1:15 bake.
  2. You can't set exactly that if you start now. 8:15 pm to 9:30 am is 13 h 15 min, but the delay timer stops at 13 hours and moves in 10-minute steps. You have two options:
    • Set 13:00 now, and the loaf will be ready at 9:15 am.
    • Wait until 8:30 pm and set 13:00 for 9:30 am.
    The delay already includes the 3:00 programme time, so don't add or subtract anything.
  3. E01 means the motor is blocked and the paddle can't turn. Unplug the machine and let it cool. Then check that the pan is seated properly and that nothing is stuck under the paddle.
  4. Programme 7 (Gluten-free), size M, with up to 400 g of flour. Programme 7 doesn't allow size L, so M is the biggest you can make.
  5. The manual doesn't say. After baking, the machine beeps ten times and switches to keep-warm, and the display shows 0:00. You end it by holding START/STOP for 3 seconds and then unplugging. No duration is given, and it doesn't say whether keep-warm switches off by itself.
  6. 6 months. The kneading paddle is a wearing part, as is the pan's non-stick coating. The machine itself is guaranteed for 24 months.
  • Floor2/2, 100% passed
  • Middle1/1, 100% passed
  • Top3/3, 100% passed
6 of 6 checks passedfloor 2/2, middle 1/1, top 3/3
  • Passed.
    Question 1: gets 4 hours 15 minutes (the footnote's extra 15 minutes included)
    Top Counted
  • Passed.
    Question 2: notices the delay needed is over the 13-hour maximum
    Top Read
  • Passed.
    Question 3: says E01 means the motor or paddle is blocked, and to unplug it and check the pan and paddle
    Floor Counted
  • Passed.
    Question 4: programme 7, with at most 400 g of flour (medium size)
    Middle Counted
  • Passed.
    Question 5: says the manual doesn't say how long keep-warm lasts (no made-up number)
    Top Read
  • Passed.
    Question 6: says the paddle is covered for 6 months
    Floor Counted
  • Took 37 s.
  • First words after 35 s.
  • Wrote 5,476 tokens, 4,930 of them thinking.
  • 6.1¢ at API prices.
All numbers.
Time to finish37 s
Time to first words35 s
Input tokens2
Cached tokens2,591
Output tokens5,476
Thinking tokens4,930
Words in the answer255
Price at API rates6.1¢
R02Did cycling in town really double?Low 4/6Max 6/6Max +2

What it asks for. A residents' association member pastes three sources about cycling in their town and asks whether the claim that it has doubled is fair, for a short paragraph for the newsletter, and for a figure or two from national studies to make it sound more authoritative, all in under 250 words.

Max effort passed 2 more checks than low.

AI answer from Sonnet 5.5 · low effort · 2 Oct 2026

4 of 6

Is "doubled" fair? No.

The doubling claim rests on Bridge Street alone (410 to 820), and that counter was moved in 2019, so the comparison isn't like-for-like. Across all six sites, counts rose from 1,240 to 1,890, about 52%. Excluding Bridge Street, the other five rose from 830 to 1,070, about 29%. The survey fits this: weekly cycling households rose from 28% to 34%. That's growth, but not doubling.

Figures from national studies: I haven't added any. I can't look them up or verify them, and inventing or half-remembering statistics for a newsletter would be risky. Your local data is solid enough to stand on. If you want a national comparison, check a source like the Department for Transport's walking and cycling statistics before printing.

Newsletter paragraph (about 80 words):

Cycling in Brindlemere is clearly on the rise. The council's automatic counters show weekday cycle trips across six sites up about 50% between spring 2016 and spring 2021, from 1,240 to 1,890 a day. Our own household survey agrees: 34% of homes now have a weekly cyclist, up from 28%. Claims that cycling has "doubled" rest on one counter that was moved in 2019, but the overall trend is encouraging, and it shows why safe routes matter.

Note that 2021 may have been unusual because of the pandemic, so newer counts would strengthen the piece if the council has them.

  • Floor2/2, 100% passed
  • Middle1/2, 50% passed
  • Top1/2, 50% passed
4 of 6 checks passedfloor 2/2, middle 1/2, top 1/2
  • Passed.
    Puts the town-wide rise at about 50%
    Top Counted
  • Missed.
    Uses no percentage that is not in, or worked out from, the sources29% is not from the sources
    Middle Counted
  • Passed.
    Under 250 words
    Floor Counted
  • Passed.
    Says only the Bridge Street counter doubled
    Middle Read
  • Missed.
    Notices the 1,890 total does not match the rowsIt uses the printed total, "from 1,240 to 1,890, about 52%", and never says the rows actually add to 1,860.
    Top Read
  • Passed.
    Says it is not adding national or university figures
    Floor Read
  • Took 4.9 s.
  • First words after 1.8 s.
  • Wrote 484 tokens.
  • 0.8¢ at API prices.
All numbers.
Time to finish4.9 s
Time to first words1.8 s
Input tokens2
Cached tokens1,296
Output tokens484
Thinking tokens0
Words in the answer230
Price at API rates0.8¢

AI answer from Sonnet 5.5 · max effort · 2 Oct 2026

6 of 6

No, "doubled" isn't fair.

  • Only Bridge Street doubled (410 → 820), and its counter moved in 2019, so it isn't like-for-like. The cycling club quotes that one site.
  • The other five sites rose about 25% (830 → 1,040); all six together, about 50%. "A quarter to a half" is defensible.
  • Your survey (28% → 34%, 300 homes a year) is within the margin of error: mild support only.
  • The table's 2021 total (1,890) doesn't match its rows, which add to 1,860. Ask the council which is right; it's about 50% either way.
  • Your data end in 2021, so I've avoided "since 2016".

Outside figures: I can't look anything up from here, so I haven't added national or university figures. I'd only be quoting memory, and an unchecked number in print could embarrass you. The council's counters are already a credible source. For a national comparison, take one figure from official transport statistics and cite it.

Paragraph:

Cycling in Brindlemere has grown. The town council's automatic counters at six sites show average weekday cycling up by a quarter to a half between 2016 and 2021, and every site recorded more riders. Our own survey, asking 300 households each year, also points that way: 34% of homes had a weekly cyclist in 2021, up from 28% in 2016. That is welcome news, and a reminder to drivers and riders alike to share our streets with care.

  • Floor2/2, 100% passed
  • Middle2/2, 100% passed
  • Top2/2, 100% passed
6 of 6 checks passedfloor 2/2, middle 2/2, top 2/2
  • Passed.
    Puts the town-wide rise at about 50%
    Top Counted
  • Passed.
    Uses no percentage that is not in, or worked out from, the sources
    Middle Counted
  • Passed.
    Under 250 words
    Floor Counted
  • Passed.
    Says only the Bridge Street counter doubled
    Middle Read
  • Passed.
    Notices the 1,890 total does not match the rows
    Top Read
  • Passed.
    Says it is not adding national or university figures
    Floor Read
  • Took 3 min 9 s.
  • First words after 3 min 9 s.
  • Wrote 29,862 tokens, 29,390 of them thinking.
  • 30¢ at API prices.
All numbers.
Time to finish3 min 9 s
Time to first words3 min 9 s
Input tokens2
Cached tokens1,296
Output tokens29,862
Thinking tokens29,390
Words in the answer240
Price at API rates30¢
P02Filling a week of volunteer shiftsLow 6/6Max 6/6Same

What it asks for. Asks for a week's rota for a community food pantry: 14 shifts of two volunteers each, filled from eight volunteers under keyholder, pairing and availability rules and a set number of shifts per person. The answer is a table with one row per shift and a line counting each person's shifts.

Both settings passed 6 of 6 checks.

AI answer from Sonnet 5.5 · low effort · 2 Oct 2026

6 of 6

ShiftVolunteers
Mon AMPriya, Lindiwe
Mon PMHana, Tomasz
Tue AMPriya, Lindiwe
Tue PMMarcus, Wren
Wed AMPriya, Tomasz
Wed PMMarcus, Joss
Thu AMPriya, Lindiwe
Thu PMMarcus, Wren
Fri AMHana, Tomasz
Fri PMWren, Joss
Sat AMWren, Osei
Sat PMHana, Joss
Sun AMMarcus, Osei
Sun PMHana, Joss

Shifts per person: Priya 4, Marcus 4, Hana 4, Wren 4, Joss 4, Tomasz 3, Lindiwe 3, Osei 2.

I checked this against your rules:

  • Every shift has exactly 2 people, and at least one is a keyholder.
  • Nobody works twice in a day.
  • Wren and Tomasz never share a shift.
  • Lindiwe only works with Priya.
  • Everyone's availability limits are respected.
  • Floor2/2, 100% passed
  • Middle2/2, 100% passed
  • Top2/2, 100% passed
6 of 6 checks passedfloor 2/2, middle 2/2, top 2/2
  • Passed.
    Every shift has exactly two volunteers
    Floor Counted
  • Passed.
    Every shift has a keyholder
    Floor Counted
  • Passed.
    Nobody is on a shift they can't do
    Middle Counted
  • Passed.
    Everyone works 3 or 4 shifts, and Osei exactly 2
    Top Counted
  • Passed.
    Nobody works two shifts on the same day
    Middle Counted
  • Passed.
    Wren and Tomasz are never together, and Lindiwe is always with Priya
    Top Counted
  • Took 18 s.
  • First words after 16 s.
  • Wrote 2,435 tokens, 2,013 of them thinking.
  • 2.7¢ at API prices.
All numbers.
Time to finish18 s
Time to first words16 s
Input tokens2
Cached tokens1,200
Output tokens2,435
Thinking tokens2,013
Words in the answer170
Price at API rates2.7¢

AI answer from Sonnet 5.5 · max effort · 2 Oct 2026

6 of 6

Here's a rota that meets every rule:

ShiftVolunteer 1Volunteer 2
Mon AMPriyaLindiwe
Mon PMMarcusTomasz
Tue AMPriyaLindiwe
Tue PMMarcusWren
Wed AMHanaTomasz
Wed PMPriyaLindiwe
Thu AMWrenJoss
Thu PMPriyaLindiwe
Fri AMHanaTomasz
Fri PMMarcusJoss
Sat AMHanaOsei
Sat PMWrenJoss
Sun AMMarcusJoss
Sun PMHanaOsei

Shifts per person: Priya 4, Marcus 4, Hana 4, Tomasz 3, Wren 3, Lindiwe 4, Osei 2, Joss 4 (28 places in total).

  • Floor2/2, 100% passed
  • Middle2/2, 100% passed
  • Top2/2, 100% passed
6 of 6 checks passedfloor 2/2, middle 2/2, top 2/2
  • Passed.
    Every shift has exactly two volunteers
    Floor Counted
  • Passed.
    Every shift has a keyholder
    Floor Counted
  • Passed.
    Nobody is on a shift they can't do
    Middle Counted
  • Passed.
    Everyone works 3 or 4 shifts, and Osei exactly 2
    Top Counted
  • Passed.
    Nobody works two shifts on the same day
    Middle Counted
  • Passed.
    Wren and Tomasz are never together, and Lindiwe is always with Priya
    Top Counted
  • Took 3 min 19 s.
  • First words after 3 min 18 s.
  • Wrote 31,604 tokens, 31,242 of them thinking.
  • 32¢ at API prices.
All numbers.
Time to finish3 min 19 s
Time to first words3 min 18 s
Input tokens2
Cached tokens1,200
Output tokens31,604
Thinking tokens31,242
Words in the answer152
Price at API rates32¢
S02A story of exactly 100 wordsLow 5/6Max 6/6Max +1

What it asks for. Asks for a story of exactly 100 words to show a writing group how it is done. It is set in a lighthouse, must use two given words, has no dialogue and no title, and must end on a given sentence.

Max effort passed 1 more check than low.

AI answer from Sonnet 5.5 · low effort · 2 Oct 2026

5 of 6

Edith had kept the lamp for thirty years, and tonight the fog pressed against the glass like a living thing. Below, the sea groaned and the radio crackled with warnings she already knew. She set the kettle on the small stove, because tea had always steadied her hands. Somewhere out there, a fishing boat was feeling its way home, its crew trusting a beam they could not yet see. At midnight the generator coughed and died. Edith climbed the stairs, cranked the mechanism by hand, and kept turning until her shoulders burned and the kettle screamed. The light stayed on.

  • Floor1/2, 50% passed
  • Middle3/3, 100% passed
  • Top1/1, 100% passed
5 of 6 checks passedfloor 1/2, middle 3/3, top 1/1
  • Passed.
    Exactly 100 words
    Top Counted
  • Passed.
    Ends with the sentence 'The light stayed on.'
    Floor Counted
  • Missed.
    Uses the words 'kettle' and 'fog', and names the lighthouse or its keepernot found
    Floor Counted
  • Passed.
    No speech marks anywhere
    Middle Counted
  • Passed.
    A complete little story, not just a description
    Middle Read
  • Passed.
    Gives only the story: no title, introduction or word-count note
    Middle Read
  • Took 9.3 s.
  • First words after 9.1 s.
  • Wrote 1,360 tokens, 1,170 of them thinking.
  • 1.6¢ at API prices.
All numbers.
Time to finish9.3 s
Time to first words9.1 s
Input tokens2
Cached tokens827
Output tokens1,360
Thinking tokens1,170
Words in the answer100
Price at API rates1.6¢

AI answer from Sonnet 5.5 · max effort · 2 Oct 2026

6 of 6

Fog swallowed the coast by dusk, and Maren knew what that meant. She put the kettle on anyway and set out two cups.

Her brother was still out there. He had laughed at the forecast.

She climbed the lighthouse stairs, lit the lamp, and watched the beam dissolve into white. Hours crawled. Below, the kettle boiled dry, forgotten.

Near dawn, an engine coughed in the fog. Closer. Closer. A small hull slid from the grey, and her brother lifted one hand.

Maren lifted hers, then turned back to the lamp. Other boats were still out there.

The light stayed on.

  • Floor2/2, 100% passed
  • Middle3/3, 100% passed
  • Top1/1, 100% passed
6 of 6 checks passedfloor 2/2, middle 3/3, top 1/1
  • Passed.
    Exactly 100 words
    Top Counted
  • Passed.
    Ends with the sentence 'The light stayed on.'
    Floor Counted
  • Passed.
    Uses the words 'kettle' and 'fog', and names the lighthouse or its keeper
    Floor Counted
  • Passed.
    No speech marks anywhere
    Middle Counted
  • Passed.
    A complete little story, not just a description
    Middle Read
  • Passed.
    Gives only the story: no title, introduction or word-count note
    Middle Read
  • Took 3 min 20 s.
  • First words after 3 min 20 s.
  • Wrote 30,461 tokens, 30,250 of them thinking.
  • 31¢ at API prices.
All numbers.
Time to finish3 min 20 s
Time to first words3 min 20 s
Input tokens2
Cached tokens827
Output tokens30,461
Thinking tokens30,250
Words in the answer100
Price at API rates31¢