ModelLineup

Every answer on this page was written by an AI model, named on the answer. Answers can be wrong, out of date or made up, even when they sound sure. Check anything that matters before you rely on it.

Max effort results for this model are being collected again. We capped every answer at 32,000 output tokens, below this model’s own output limit, so many max effort answers ran out of room while still thinking and scored zero. Low effort results are not affected. The new max effort answers will replace these.

Models Opus 5.5

Claude Opus 5.5

Released
Tested
Settings
Low effort and Max effort

Opus 5.5

24 prompts, every check counted

Low effort

90

  • Floor46/49, 94% passed
  • Middle51/58, 88% passed
  • Top60/65, 92% passed

Median answer 16 s, 4.1¢.

Max effort

49

  • Floor28/49, 57% passed
  • Middle28/58, 48% passed
  • Top32/65, 49% passed

Median answer 4 min 0 s, 55¢.

Max passed 69 fewer checks, cost 8.1 times as much and took 9.7 times as long. Costs are at Anthropic’s API prices on .

What the results show

Written by an AI model from these results; every figure in it is filled in by a program. How it was made.

Opus 5.5 answered all 24 prompts at both efforts on 2 October 2026, scoring 90 at low effort and 49 at max. At max effort 11 answers ran out of room, with the output limit spent mostly on thinking.

Best at

  • Code: it passed every check on the paging bug, the search box race, the early renewals and the customers who left and came back at both efforts, and on the contacts import and the slow script at low effort. F01 F02 F03 F05 F06 F07
  • Web pages at low effort, 90 for the group: the landing page for the invoicing app and the garden page fixes passed every page test, and the dashboard from the mockup was judged a faithful build. B01 B05 B02
  • Careful reading at both efforts: it found the wrong line on the market receipt, said it did not know the invented book, and met every check on the bread maker manual questions. I01 H02 K02
  • Everyday writing and the rota: the text to the plumber, the update from a messy thread and the volunteer rota met every check at both efforts; the everyday group scored 96 at low and 98 at max. W01 W02 P02

Stumbled on

  • At max effort all the web page answers ran out of room before any page was written; the landing page, dashboard, kanban board, Minesweeper game and garden page fixes left nothing for the page tests to load. B01 B02 B03 B04 B05
  • Also at max effort the contacts import, slow script, booking plan, column rename plan and feature spec ran out of room and met no check; the cycling paragraph was cut off before it dealt with the outside figures. F06 F07 A01 A03 A04 R02
  • At low effort the column rename plan did not make dropping the old column its last step, the spec did not fully raise the expiry, role and seat questions, and the cycling paragraph took the printed total on trust. A03 A04 R02
  • Also at low effort the pull request review, the booking plan and the microservices answer all ran over length, and the booking plan left a working booking flow out of its first phase. F04 A01 A02

What max effort changed

  • Overall it scored lower at max effort: 3 prompts gained at max effort, 10 lost ground and 11 stayed the same, for about 8.1 times the cost and 9.7 times the time.
  • Every prompt that lost ground had run out of room: thinking took 527,544 tokens at max effort against 11,719 at low. Web pages fell from 90 to 0, planning from 76 to 25 and code from 98 to 71. B01 F06 F07 A01 A03 A04
  • Where it finished, max effort helped: the pull request review went from 6 of 7 checks to 7 of 7, the microservices answer from 7 of 8 to 8 of 8, and the story used every required word. F04 A02 S02

In short

At low effort Opus 5.5 scored level with Opus 5 at low effort, at a median 15.7 seconds and $0.041 an answer. Overall it scored lower at max effort, because 11 of its max effort answers ran out of room.

Worth knowing

  • Each prompt was answered once at each effort, so a single answer can move a prompt's score, and small gaps between scores are treated as level.
  • Every model gets the same output allowance. An answer that runs out of room is scored on whatever it did write, so an unfinished page or program usually fails its tests.
  • Costs are what the same tokens would cost at the published API prices, not what anyone paid.

Score by kind of work

Building web pages

  • Low effort
    90
  • Max effort
    0

Code

  • Low effort
    98
  • Max effort
    71

Planning and architecture

  • Low effort
    76
  • Max effort
    25

Everyday

  • Low effort
    96
  • Max effort
    98

Every answer

Each row opens to the low and max answers side by side. Each prompt is described on the prompts page.

Building web pages

B01A landing page for a small invoicing appLow 9/9Max 0/9Max −9

What it asks for. Asks for a one-page landing site for a small invoicing app for freelancers, as a single HTML file with no libraries and nothing loaded from outside. The page needs a headline, three named features, three price plans with a monthly and yearly switch (yearly gives two months free and is quoted as one total for the year), a four-question FAQ that opens and closes, and a sign-up button. It has to read well on a phone.

Max effort passed 9 fewer checks than low.

AI answer from Opus 5.5 · low effort · 2 Oct 2026

9 of 9

Screenshot of the page built in this answer, at a desktop width of 1280 pixels
Our screenshot, 1280 by 800. Open it for the full length.
Screenshot of the page built in this answer, at a phone width of 375 pixels
Phone, 375 by 812.
More screenshots (2).
Screenshot of the page built in this answer, at a desktop width of 1280 pixelsScreenshot of the page built in this answer, at a phone width of 375 pixels
Show the code and the answer.

Here's the complete file. Save it as index.html and open it in a browser.

<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Brindlebill: Simple invoicing for freelancers</title>
<meta name="description" content="Brindlebill helps freelancers send invoices, chase payments gently and get paid by card or bank.">
<style>
  :root {
    --ink: #1d1a17;
    --muted: #57514a;
    --bg: #fbf8f3;
    --card: #ffffff;
    --line: #e4ddd2;
    --brand: #8a4b1f;
    --brand-dark: #6e3a15;
    --brand-soft: #f3e6d8;
    --radius: 12px;
  }

  *, *::before, *::after { box-sizing: border-box; }

  html { -webkit-text-size-adjust: 100%; }

  body {
    margin: 0;
    font-family: system-ui, -apple-system, "Segoe UI", Roboto, Helvetica, Arial, sans-serif;
    font-size: 1rem;           /* 16px minimum */
    line-height: 1.6;
    color: var(--ink);
    background: var(--bg);
  }

  .wrap {
    width: 100%;
    max-width: 1040px;
    margin: 0 auto;
    padding: 0 1.25rem;
  }

  .skip-link {
    position: absolute;
    left: -9999px;
    top: 0;
    background: var(--ink);
    color: #fff;
    padding: 0.75rem 1rem;
    z-index: 10;
  }
  .skip-link:focus { left: 0; }

  :focus-visible {
    outline: 3px solid var(--brand);
    outline-offset: 3px;
  }

  h1, h2, h3 { line-height: 1.2; margin: 0 0 0.75rem; }
  h1 { font-size: clamp(2rem, 7vw, 3.25rem); letter-spacing: -0.02em; }
  h2 { font-size: clamp(1.6rem, 5vw, 2.25rem); letter-spacing: -0.01em; }
  h3 { font-size: 1.25rem; }
  p  { margin: 0 0 1rem; }

  /* Header */
  .site-header { padding: 1.25rem 0; }
  .logo {
    font-weight: 800;
    font-size: 1.25rem;
    color: var(--ink);
    text-decoration: none;
  }
  .logo span { color: var(--brand); }

  /* Buttons */
  .btn {
    display: inline-block;
    font: inherit;
    font-size: 1.125rem;
    font-weight: 700;
    text-decoration: none;
    text-align: center;
    color: #fff;
    background: var(--brand);
    border: 0;
    border-radius: 999px;
    padding: 0.9rem 1.75rem;
    min-height: 48px;
    cursor: pointer;
    transition: background 0.15s ease;
  }
  .btn:hover { background: var(--brand-dark); }

  /* Hero */
  .hero { padding: 3rem 0 4rem; text-align: center; }
  .hero .pitch {
    font-size: 1.25rem;
    color: var(--muted);
    max-width: 36rem;
    margin: 0 auto 2rem;
  }

  section { padding: 3.5rem 0; }
  .section-intro { color: var(--muted); max-width: 38rem; }
  .center { text-align: center; }
  .center .section-intro { margin-left: auto; margin-right: auto; }

  /* Features */
  .features { background: var(--card); border-block: 1px solid var(--line); }
  .feature-grid {
    list-style: none;
    margin: 2rem 0 0;
    padding: 0;
    display: grid;
    gap: 1.5rem;
    grid-template-columns: 1fr;
  }
  .feature-grid li {
    padding: 1.5rem;
    border-radius: var(--radius);
    background: var(--bg);
    border: 1px solid var(--line);
  }
  .feature-grid p { margin: 0; color: var(--muted); }

  /* Pricing */
  .billing-switch {
    display: inline-flex;
    margin: 1rem 0 2rem;
    padding: 4px;
    background: var(--brand-soft);
    border-radius: 999px;
  }
  .billing-switch button {
    font: inherit;
    font-size: 1rem;
    font-weight: 600;
    color: var(--ink);
    background: transparent;
    border: 0;
    border-radius: 999px;
    padding: 0.6rem 1.25rem;
    min-height: 44px;
    cursor: pointer;
  }
  .billing-switch button[aria-pressed="true"] {
    background: var(--card);
    color: var(--brand-dark);
    box-shadow: 0 1px 3px rgba(0,0,0,0.12);
  }
  .save-note { display: block; font-size: 1rem; color: var(--muted); margin-top: -1rem; margin-bottom: 2rem; }

  .plans {
    display: grid;
    gap: 1.5rem;
    grid-template-columns: 1fr;
    text-align: left;
  }
  .plan {
    display: flex;
    flex-direction: column;
    background: var(--card);
    border: 1px solid var(--line);
    border-radius: var(--radius);
    padding: 1.75rem;
  }
  .plan.featured { border: 2px solid var(--brand); }
  .price {
    font-size: 2.5rem;
    font-weight: 800;
    margin: 0.25rem 0 1.25rem;
    line-height: 1.1;
  }
  .price .period {
    font-size: 1rem;
    font-weight: 500;
    color: var(--muted);
  }
  .plan ul {
    list-style: none;
    padding: 0;
    margin: 0;
  }
  .plan li {
    padding: 0.5rem 0 0.5rem 1.75rem;
    position: relative;
    border-top: 1px solid var(--line);
  }
  .plan li::before {
    content: "✓";
    position: absolute;
    left: 0;
    color: var(--brand);
    font-weight: 700;
  }

  /* FAQ */
  .faq-list { max-width: 44rem; margin: 2rem auto 0; }
  .faq-list details {
    background: var(--card);
    border: 1px solid var(--line);
    border-radius: var(--radius);
    margin-bottom: 0.75rem;
  }
  .faq-list summary {
    list-style: none;
    cursor: pointer;
    padding: 1rem 3rem 1rem 1.25rem;
    position: relative;
    min-height: 48px;
  }
  .faq-list summary::-webkit-details-marker { display: none; }
  .faq-list summary h3 {
    display: inline;
    font-size: 1.125rem;
    margin: 0;
  }
  .faq-list summary::after {
    content: "+";
    position: absolute;
    right: 1.25rem;
    top: 50%;
    transform: translateY(-50%);
    font-size: 1.5rem;
    color: var(--brand);
  }
  .faq-list details[open] summary::after { content: "–"; }
  .faq-list details p {
    margin: 0;
    padding: 0 1.25rem 1.25rem;
    color: var(--muted);
  }

  /* Final CTA */
  .cta {
    background: var(--brand-soft);
    text-align: center;
  }

  footer {
    padding: 2rem 0;
    text-align: center;
    color: var(--muted);
  }

  @media (min-width: 720px) {
    .feature-grid { grid-template-columns: repeat(3, 1fr); }
    .plans { grid-template-columns: repeat(3, 1fr); }
    section { padding: 5rem 0; }
    .hero { padding: 5rem 0 6rem; }
  }

  @media (prefers-reduced-motion: reduce) {
    * { transition: none !important; }
  }
</style>
</head>
<body>

<a class="skip-link" href="#main">Skip to content</a>

<header class="site-header">
  <div class="wrap">
    <a class="logo" href="#">Brindle<span>bill</span></a>
  </div>
</header>

<main id="main">

  <section class="hero" aria-labelledby="hero-title">
    <div class="wrap">
      <h1 id="hero-title">Get paid without the paperwork</h1>
      <p class="pitch">Brindlebill sends your invoices, follows up for you and lets clients pay in a couple of clicks.</p>
      <a class="btn" href="#signup">Start free trial</a>
    </div>
  </section>

  <section class="features" aria-labelledby="features-title">
    <div class="wrap center">
      <h2 id="features-title">Everything you need to get paid</h2>
      <ul class="feature-grid">
        <li>
          <h3>Invoices in a minute</h3>
          <p>Pick a client, add your hours or items, and send a clean, professional invoice straight away.</p>
        </li>
        <li>
          <h3>Gentle reminders</h3>
          <p>Brindlebill nudges late payers politely on your behalf, so you never have to send the awkward email.</p>
        </li>
        <li>
          <h3>Card and bank payments</h3>
          <p>Clients can pay by card or bank transfer right from the invoice, and you see the moment it lands.</p>
        </li>
      </ul>
    </div>
  </section>

  <section id="pricing" aria-labelledby="pricing-title">
    <div class="wrap center">
      <h2 id="pricing-title">Simple pricing</h2>
      <p class="section-intro">Choose how you'd like to pay. Yearly plans get two months free.</p>

      <div class="billing-switch" role="group" aria-label="Billing period">
        <button type="button" data-period="monthly" aria-pressed="true">Monthly</button>
        <button type="button" data-period="yearly" aria-pressed="false">Yearly</button>
      </div>

      <div class="plans">
        <article class="plan" aria-labelledby="plan-starter">
          <h3 id="plan-starter">Starter</h3>
          <p class="price" aria-live="polite">
            <span class="amount" data-monthly="9" data-yearly="90">$9</span>
            <span class="period">/ month</span>
          </p>
          <ul>
            <li>5 clients</li>
            <li>PDF invoices</li>
          </ul>
        </article>

        <article class="plan featured" aria-labelledby="plan-pro">
          <h3 id="plan-pro">Pro</h3>
          <p class="price" aria-live="polite">
            <span class="amount" data-monthly="19" data-yearly="190">$19</span>
            <span class="period">/ month</span>
          </p>
          <ul>
            <li>Unlimited clients</li>
            <li>Automatic reminders</li>
            <li>Your own logo</li>
          </ul>
        </article>

        <article class="plan" aria-labelledby="plan-team">
          <h3 id="plan-team">Team</h3>
          <p class="price" aria-live="polite">
            <span class="amount" data-monthly="49" data-yearly="490">$49</span>
            <span class="period">/ month</span>
          </p>
          <ul>
            <li>Everything in Pro</li>
            <li>5 seats</li>
            <li>A shared client list</li>
          </ul>
        </article>
      </div>
    </div>
  </section>

  <section class="features" aria-labelledby="faq-title">
    <div class="wrap">
      <h2 id="faq-title" class="center">Questions</h2>
      <div class="faq-list">
        <details>
          <summary><h3>Can I cancel any time?</h3></summary>
          <p>Yes. Cancel from your account page and you won't be charged again. Nothing is locked in.</p>
        </details>
        <details>
          <summary><h3>Do my clients need an account?</h3></summary>
          <p>No. They open a link, see the invoice and pay. They never have to sign up.</p>
        </details>
        <details>
          <summary><h3>Which currencies can I invoice in?</h3></summary>
          <p>You can invoice in 30 currencies, and each invoice shows the amount in your client's currency.</p>
        </details>
        <details>
          <summary><h3>What happens when the free trial ends?</h3></summary>
          <p>You pick a plan. If you don't, your account switches to view-only and your invoices stay safe for 90 days.</p>
        </details>
      </div>
    </div>
  </section>

  <section class="cta" id="signup" aria-labelledby="signup-title">
    <div class="wrap">
      <h2 id="signup-title">Ready to send your first invoice?</h2>
      <p>Try Brindlebill free. No card needed to look around.</p>
      <a class="btn" href="#signup">Start free trial</a>
    </div>
  </section>

</main>

<footer>
  <div class="wrap">
    <p>&copy; <span id="year"></span> Brindlebill</p>
  </div>
</footer>

<script>
  (function () {
    var buttons = document.querySelectorAll('.billing-switch button');
    var amounts = document.querySelectorAll('.price .amount');

    function setPeriod(period) {
      buttons.forEach(function (b) {
        b.setAttribute('aria-pressed', String(b.dataset.period === period));
      });
      amounts.forEach(function (el) {
        el.textContent = '$' + el.dataset[period];
        el.nextElementSibling.textContent = period === 'yearly' ? '/ year' : '/ month';
      });
    }

    buttons.forEach(function (b) {
      b.addEventListener('click', function () { setPeriod(b.dataset.period); });
    });

    document.getElementById('year').textContent = new Date().getFullYear();
  })();
</script>

</body>
</html>

Notes

  • Yearly prices are 10 × the monthly price: Starter $90, Pro $190, Team $490, each shown as "/ year".
  • FAQ uses the built-in <details>/<summary> elements. They open and close without JavaScript and work with a keyboard and screen readers.
  • Headings: one h1, an h2 for each section, and an h3 for each feature, plan and question.
  • Mobile: all text is at least 16px. Columns stack on narrow screens, and buttons are at least 44px tall so they're easy to tap.
  • Sign-up button: both "Start free trial" buttons currently point to #signup. Change the href to your real sign-up URL. The line "No card needed to look around" under the bottom button is a placeholder, so edit it to match how your trial actually works.
  • I wrote the headline, pitch and feature sentences myself, so change them as you like.
  • Floor2/2, 100% passed
  • Middle4/4, 100% passed
  • Top3/3, 100% passed
9 of 9 checks passedfloor 2/2, middle 4/4, top 3/3
  • Passed.
    The page loads with no errors in the console
    Floor Page test
  • Passed.
    Every section is there, with the exact plan names, prices and text
    Floor Page test
  • Passed.
    No sideways scroll on a 375 px phone
    Middle Page test
  • Passed.
    Nothing is loaded from outside the file
    Middle Page test
  • Passed.
    One h1, no skipped heading levels, text 16 px or larger, enough contrast
    Middle Page test
  • Passed.
    The page looks finished: clear hierarchy, pricing easy to compare, nothing broken
    Middle Read
  • Passed.
    Yearly shows $90, $190 and $490, and Monthly brings back $9, $19, $49
    Top Page test
  • Passed.
    Each FAQ question opens and closes from the keyboard
    Top Page test
  • Passed.
    Still no sideways scroll on a phone with Yearly on and every FAQ open
    Top Page test
  • Took 42 s.
  • First words after 4.7 s.
  • Wrote 5,389 tokens, 209 of them thinking.
  • 11¢ at API prices.
All numbers.
Time to finish42 s
Time to first words4.7 s
Input tokens2
Cached tokens1,272
Output tokens5,389
Thinking tokens209
Words in the answer1,314
Price at API rates11¢

AI answer from Opus 5.5 · max effort · 2 Oct 2026

0 of 9

The answer was cut off at the token limit.

  • Floor0/2, 0% passed
  • Middle0/4, 0% passed
  • Top0/3, 0% passed
0 of 9 checks passedfloor 0/2, middle 0/4, top 0/3
  • Missed.
    The page loads with no errors in the consoleNo answer
    Floor Page test
  • Missed.
    Every section is there, with the exact plan names, prices and textNo answer
    Floor Page test
  • Missed.
    No sideways scroll on a 375 px phoneNo answer
    Middle Page test
  • Missed.
    Nothing is loaded from outside the fileNo answer
    Middle Page test
  • Missed.
    One h1, no skipped heading levels, text 16 px or larger, enough contrastNo answer
    Middle Page test
  • Missed.
    The page looks finished: clear hierarchy, pricing easy to compare, nothing brokenNo answer
    Middle Read
  • Missed.
    Yearly shows $90, $190 and $490, and Monthly brings back $9, $19, $49No answer
    Top Page test
  • Missed.
    Each FAQ question opens and closes from the keyboardNo answer
    Top Page test
  • Missed.
    Still no sideways scroll on a phone with Yearly on and every FAQ openNo answer
    Top Page test
  • Took 5 min 23 s.
  • Wrote 32,000 tokens, 32,000 of them thinking.
  • 65¢ at API prices.
All numbers.
Time to finish5 min 23 s
Time to first wordsnot recorded
Input tokens2
Cached tokens1,272
Output tokens32,000
Thinking tokens32,000
Words in the answer0
Price at API rates65¢
B02Build a dashboard from a picture of its designLow 6/7Max 0/7Max −6

What it asks for. Sends a picture of a designer's mockup, an admin dashboard for a small bike repair shop, and asks for it to be built as one HTML file that matches the layout, colours, spacing and text as closely as it can. The bar chart has to be drawn to scale. On a phone the page must stack into one column, hide the sidebar behind a menu button and never scroll sideways.

Max effort passed 6 fewer checks than low.

AI answer from Opus 5.5 · low effort · 2 Oct 2026

6 of 7

Screenshot of the page built in this answer, at a desktop width of 1280 pixels
Our screenshot, 1280 by 800. Open it for the full length.
Screenshot of the page built in this answer, at a phone width of 375 pixels
Phone, 375 by 812.
Show the code and the answer.

Here's the complete file. Save it as something like index.html and open it in a browser.

<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Overview · Tarnwick Bikeworks</title>
<style>
  :root {
    --sidebar: #1c2a43;
    --sidebar-line: #2b3a54;
    --sidebar-text: #c9d2e0;
    --blue: #2f6fea;
    --bg: #f2f4f8;
    --card: #ffffff;
    --border: #e3e7ee;
    --text: #172033;
    --muted: #5f6b7d;
    --green: #15803d;
    --red: #cf3a24;
    --unit: 4px; /* chart scale: 1 job = 4px of bar height */
  }

  * { box-sizing: border-box; }

  html, body {
    margin: 0;
    padding: 0;
  }

  body {
    font-family: -apple-system, BlinkMacSystemFont, "Segoe UI", Roboto, "Helvetica Neue", Arial, sans-serif;
    background: var(--bg);
    color: var(--text);
    font-size: 14px;
    -webkit-font-smoothing: antialiased;
    overflow-x: hidden;
  }

  /* ---------- Sidebar ---------- */
  .sidebar {
    position: fixed;
    top: 0; left: 0; bottom: 0;
    width: 240px;
    background: var(--sidebar);
    color: #fff;
    z-index: 30;
  }
  .brand {
    height: 65px;
    display: flex;
    align-items: center;
    padding: 0 24px;
    font-weight: 700;
    font-size: 16px;
    letter-spacing: -0.1px;
    border-bottom: 1px solid var(--sidebar-line);
  }
  .nav {
    list-style: none;
    margin: 0;
    padding: 23px 12px;
  }
  .nav li + li { margin-top: 4px; }
  .nav a {
    display: flex;
    align-items: center;
    gap: 12px;
    height: 40px;
    padding: 0 14px;
    border-radius: 6px;
    color: var(--sidebar-text);
    text-decoration: none;
    font-size: 15px;
  }
  .nav a:hover { background: rgba(255,255,255,0.06); }
  .nav a.active {
    background: var(--blue);
    color: #fff;
    font-weight: 600;
  }
  .nav .icon {
    width: 17px;
    height: 17px;
    border: 1.6px solid currentColor;
    border-radius: 4px;
    flex: none;
  }

  /* ---------- Main ---------- */
  .main {
    margin-left: 240px;
    min-height: 100vh;
    min-width: 0;
  }

  .topbar {
    height: 65px;
    background: #fff;
    border-bottom: 1px solid var(--border);
    display: flex;
    align-items: center;
    padding: 0 32px;
    gap: 16px;
  }
  .topbar h1 {
    margin: 0;
    font-size: 22px;
    font-weight: 700;
    letter-spacing: -0.3px;
    flex: 1;
    min-width: 0;
  }
  .menu-btn {
    display: none;
    width: 38px;
    height: 38px;
    border: 1px solid var(--border);
    background: #fff;
    border-radius: 8px;
    padding: 0;
    cursor: pointer;
    align-items: center;
    justify-content: center;
    flex: none;
  }
  .menu-btn span,
  .menu-btn span::before,
  .menu-btn span::after {
    display: block;
    width: 18px;
    height: 2px;
    background: var(--text);
    border-radius: 2px;
    position: relative;
  }
  .menu-btn span::before,
  .menu-btn span::after {
    content: "";
    position: absolute;
    left: 0;
  }
  .menu-btn span::before { top: -6px; }
  .menu-btn span::after  { top: 6px; }

  .search {
    width: 292px;
    height: 36px;
    display: flex;
    align-items: center;
    gap: 9px;
    padding: 0 12px;
    background: #f1f4f8;
    border: 1px solid #dde2ea;
    border-radius: 6px;
    margin-right: 30px;
  }
  .search svg { flex: none; }
  .search input {
    border: 0;
    background: transparent;
    outline: none;
    font: inherit;
    font-size: 14px;
    color: var(--text);
    width: 100%;
    min-width: 0;
  }
  .search input::placeholder { color: #5f6b7d; }
  .avatar {
    width: 36px;
    height: 36px;
    border-radius: 50%;
    background: var(--sidebar);
    color: #fff;
    font-size: 12px;
    font-weight: 700;
    display: flex;
    align-items: center;
    justify-content: center;
    flex: none;
  }

  .content {
    padding: 31px 32px 48px;
  }

  .card {
    background: var(--card);
    border: 1px solid var(--border);
    border-radius: 10px;
  }

  /* ---------- Stats ---------- */
  .stats {
    display: grid;
    grid-template-columns: repeat(3, 1fr);
    gap: 32px;
  }
  .stat {
    padding: 18px 20px 10px;
    min-height: 109px;
  }
  .stat .label {
    font-size: 13px;
    color: #4b5668;
    font-weight: 500;
  }
  .stat .value {
    font-size: 30px;
    font-weight: 700;
    letter-spacing: -0.5px;
    margin: 4px 0 6px;
    line-height: 1.2;
  }
  .delta {
    font-size: 13px;
    font-weight: 600;
    display: flex;
    align-items: center;
    gap: 6px;
  }
  .delta.up { color: var(--green); }
  .delta.down { color: var(--red); }
  .delta .tri {
    width: 0; height: 0;
    border-left: 5px solid transparent;
    border-right: 5px solid transparent;
  }
  .delta.up .tri { border-bottom: 8px solid currentColor; }
  .delta.down .tri { border-top: 8px solid currentColor; }

  /* ---------- Chart ---------- */
  .chart-card {
    margin-top: 32px;
    padding: 18px 23px 9px;
  }
  .card-head {
    display: flex;
    justify-content: space-between;
    align-items: baseline;
  }
  .card-head h2 {
    margin: 0;
    font-size: 15px;
    font-weight: 700;
  }
  .card-head .range {
    font-size: 13px;
    color: var(--muted);
  }
  .chart {
    position: relative;
    margin-top: 22px;
    height: 148px;              /* 124px for the tallest bar + room for its label */
    border-bottom: 1px solid #dfe3ea;
    display: flex;
  }
  .chart .grid {
    position: absolute;
    left: 0; right: 0;
    bottom: calc(15.5 * var(--unit));
    border-top: 1px solid #edf0f4;
  }
  .col {
    flex: 1;
    display: flex;
    flex-direction: column;
    justify-content: flex-end;
    align-items: center;
    position: relative;
  }
  .bar {
    width: 56px;
    max-width: 42%;
    height: calc(var(--v) * var(--unit));
    background: var(--blue);
    border-radius: 3px 3px 0 0;
  }
  .bar-val {
    font-size: 13px;
    font-weight: 700;
    margin-bottom: 6px;
  }
  .days {
    display: flex;
  }
  .days span {
    flex: 1;
    text-align: center;
    font-size: 13px;
    color: var(--muted);
    padding-top: 8px;
  }

  /* ---------- Table ---------- */
  .jobs-card {
    margin-top: 32px;
    padding-top: 17px;
    padding-bottom: 12px;
    overflow: hidden;
  }
  .jobs-card h2 {
    margin: 0 0 12px;
    padding: 0 23px;
    font-size: 15px;
    font-weight: 700;
  }
  .table-wrap {
    overflow-x: auto;
    -webkit-overflow-scrolling: touch;
  }
  table {
    width: 100%;
    border-collapse: collapse;
    min-width: 640px;
  }
  thead th {
    background: #f6f8fb;
    text-align: left;
    font-size: 12px;
    font-weight: 600;
    color: var(--muted);
    height: 32px;
    padding: 0 12px;
  }
  td {
    height: 44px;
    padding: 0 12px;
    font-size: 14px;
    border-top: 1px solid #edf0f4;
    white-space: nowrap;
  }
  tbody tr:first-child td { border-top: 0; }
  th:first-child, td:first-child { padding-left: 23px; width: 96px; }
  th:nth-child(2) { width: 178px; }
  th:nth-child(3) { width: 230px; }
  th:nth-child(4) { width: 120px; }
  td:first-child { font-weight: 700; }
  .pill {
    display: inline-block;
    padding: 5px 15px;
    border-radius: 999px;
    font-size: 12.5px;
    font-weight: 600;
  }
  .pill.ready   { background: #dcf3e3; color: #17663a; }
  .pill.prog    { background: #fcefc9; color: #8a5300; }
  .pill.parts   { background: #e8e3fc; color: #5536bf; }

  .overlay { display: none; }

  /* ---------- Mobile ---------- */
  @media (max-width: 860px) {
    .sidebar {
      transform: translateX(-100%);
      transition: transform 0.25s ease;
      box-shadow: none;
    }
    body.nav-open .sidebar {
      transform: translateX(0);
      box-shadow: 0 0 40px rgba(0,0,0,0.3);
    }
    .overlay {
      display: block;
      position: fixed;
      inset: 0;
      background: rgba(15, 23, 42, 0.45);
      opacity: 0;
      pointer-events: none;
      transition: opacity 0.25s ease;
      z-index: 20;
    }
    body.nav-open .overlay {
      opacity: 1;
      pointer-events: auto;
    }
    body.nav-open { overflow: hidden; }

    .main { margin-left: 0; }
    .menu-btn { display: inline-flex; }

    .topbar {
      height: auto;
      flex-wrap: wrap;
      padding: 12px 16px;
      gap: 12px;
    }
    .search {
      order: 3;
      width: 100%;
      margin-right: 0;
    }

    .content { padding: 16px 16px 32px; }
    .stats { grid-template-columns: 1fr; gap: 16px; }
    .chart-card, .jobs-card { margin-top: 16px; }
    .chart-card { padding: 16px 14px 8px; }
    .bar { max-width: 60%; }
    .days span { font-size: 12px; }
    .jobs-card h2 { padding: 0 16px; }
    th:first-child, td:first-child { padding-left: 16px; }
  }
</style>
</head>
<body>

<aside class="sidebar" id="sidebar" aria-label="Main navigation">
  <div class="brand">Tarnwick Bikeworks</div>
  <ul class="nav">
    <li><a href="#" class="active" aria-current="page"><span class="icon"></span>Overview</a></li>
    <li><a href="#"><span class="icon"></span>Jobs</a></li>
    <li><a href="#"><span class="icon"></span>Bookings</a></li>
    <li><a href="#"><span class="icon"></span>Stock</a></li>
    <li><a href="#"><span class="icon"></span>Reports</a></li>
  </ul>
</aside>
<div class="overlay" id="overlay"></div>

<div class="main">
  <header class="topbar">
    <button class="menu-btn" id="menuBtn" aria-label="Open menu" aria-controls="sidebar" aria-expanded="false"><span></span></button>
    <h1>Overview</h1>
    <label class="search">
      <svg width="16" height="16" viewBox="0 0 16 16" fill="none" aria-hidden="true">
        <circle cx="7" cy="7" r="5.25" stroke="#5f6b7d" stroke-width="1.6"/>
        <path d="M11 11l3.5 3.5" stroke="#5f6b7d" stroke-width="1.6" stroke-linecap="round"/>
      </svg>
      <input type="search" placeholder="Search jobs and bikes" aria-label="Search jobs and bikes">
    </label>
    <div class="avatar" aria-label="Account">TB</div>
  </header>

  <main class="content">
    <section class="stats">
      <div class="card stat">
        <div class="label">Revenue</div>
        <div class="value">$8,460</div>
        <div class="delta up"><span class="tri"></span>12.4% vs last week</div>
      </div>
      <div class="card stat">
        <div class="label">Jobs completed</div>
        <div class="value">130</div>
        <div class="delta up"><span class="tri"></span>8.3% vs last week</div>
      </div>
      <div class="card stat">
        <div class="label">Open bookings</div>
        <div class="value">24</div>
        <div class="delta down"><span class="tri"></span>6.0% vs last week</div>
      </div>
    </section>

    <section class="card chart-card">
      <div class="card-head">
        <h2>Completed per day</h2>
        <span class="range">Last 7 days</span>
      </div>
      <div class="chart" role="img" aria-label="Jobs completed per day: Mon 14, Tue 22, Wed 18, Thu 27, Fri 31, Sat 12, Sun 6">
        <div class="grid"></div>
        <div class="col"><span class="bar-val">14</span><div class="bar" style="--v:14"></div></div>
        <div class="col"><span class="bar-val">22</span><div class="bar" style="--v:22"></div></div>
        <div class="col"><span class="bar-val">18</span><div class="bar" style="--v:18"></div></div>
        <div class="col"><span class="bar-val">27</span><div class="bar" style="--v:27"></div></div>
        <div class="col"><span class="bar-val">31</span><div class="bar" style="--v:31"></div></div>
        <div class="col"><span class="bar-val">12</span><div class="bar" style="--v:12"></div></div>
        <div class="col"><span class="bar-val">6</span><div class="bar" style="--v:6"></div></div>
      </div>
      <div class="days">
        <span>Mon</span><span>Tue</span><span>Wed</span><span>Thu</span><span>Fri</span><span>Sat</span><span>Sun</span>
      </div>
    </section>

    <section class="card jobs-card">
      <h2>Recent jobs</h2>
      <div class="table-wrap">
        <table>
          <thead>
            <tr><th>Job</th><th>Bike</th><th>Work</th><th>Due</th><th>Status</th></tr>
          </thead>
          <tbody>
            <tr><td>#2041</td><td>Trail hardtail</td><td>Gear tune and chain</td><td>12 Mar</td><td><span class="pill ready">Ready for pickup</span></td></tr>
            <tr><td>#2040</td><td>Commuter 700c</td><td>Brake pads, both wheels</td><td>12 Mar</td><td><span class="pill prog">In progress</span></td></tr>
            <tr><td>#2039</td><td>Folding bike</td><td>New rear tyre</td><td>13 Mar</td><td><span class="pill parts">Waiting for parts</span></td></tr>
            <tr><td>#2038</td><td>Road racer</td><td>Full service</td><td>14 Mar</td><td><span class="pill prog">In progress</span></td></tr>
            <tr><td>#2037</td><td>Cargo trike</td><td>Wheel true, front</td><td>14 Mar</td><td><span class="pill ready">Ready for pickup</span></td></tr>
          </tbody>
        </table>
      </div>
    </section>
  </main>
</div>

<script>
  (function () {
    var body = document.body;
    var btn = document.getElementById('menuBtn');
    var overlay = document.getElementById('overlay');

    function setOpen(open) {
      body.classList.toggle('nav-open', open);
      btn.setAttribute('aria-expanded', open ? 'true' : 'false');
      btn.setAttribute('aria-label', open ? 'Close menu' : 'Open menu');
    }

    btn.addEventListener('click', function () {
      setOpen(!body.classList.contains('nav-open'));
    });
    overlay.addEventListener('click', function () { setOpen(false); });
    document.addEventListener('keydown', function (e) {
      if (e.key === 'Escape') setOpen(false);
    });
    window.addEventListener('resize', function () {
      if (window.innerWidth > 860) setOpen(false);
    });
  })();
</script>
</body>
</html>

How it works

  • Chart scale: Each bar's height is value × 4px, set with a CSS variable (style="--v:27"). The tallest bar, 31, comes out at 124px, which matches the mockup. To change the scale, edit --unit in one place. The faint gridline sits at the same height as in the design.
  • Phone layout (860px and narrower):
    • The sidebar slides off-screen.
    • A menu button appears next to the title. Tapping it slides the sidebar in over a dimmed backdrop.
    • Tapping the backdrop or pressing Esc closes it.
    • The stat cards stack into one column, and the search box moves to its own full-width row.
  • No sideways page scroll: The table is the only element too wide for a phone. It sits in its own scroll container, so only the table swipes sideways and the page stays put.
  • Fonts: Only the system font stack is used. On a Mac this gives the same San Francisco look as the mockup.
  • Floor2/2, 100% passed
  • Middle2/2, 100% passed
  • Top2/3, 67% passed
6 of 7 checks passedfloor 2/2, middle 2/2, top 2/3
  • Passed.
    Page loads with no console errors and nothing fetched from outside
    Floor Page test
  • Passed.
    Every word and number in the picture is on the page
    Floor Page test
  • Passed.
    Wide screen: full-height left sidebar, three cards in a row, table cells in the right columns
    Middle Page test
  • Passed.
    Sidebar, active item, accent blue and the three status pills match the picture
    Middle Page test
  • Passed.
    Chart bars run Mon to Sun with heights in proportion to their values
    Top Page test
  • Missed.
    On a phone: no sideways scroll, cards stacked, sidebar hidden until a menu button opens ittest group failed
    Top Page test
  • Passed.
    Side by side with the picture, a designer would accept it as a faithful build
    Top Read
  • Took 54 s.
  • First words after 8.5 s.
  • Wrote 6,998 tokens, 544 of them thinking.
  • 15¢ at API prices.
All numbers.
Time to finish54 s
Time to first words8.5 s
Input tokens2
Cached tokens2,327
Output tokens6,998
Thinking tokens544
Words in the answer1,524
Price at API rates15¢

AI answer from Opus 5.5 · max effort · 2 Oct 2026

0 of 7

The answer was cut off at the token limit.

  • Floor0/2, 0% passed
  • Middle0/2, 0% passed
  • Top0/3, 0% passed
0 of 7 checks passedfloor 0/2, middle 0/2, top 0/3
  • Missed.
    Page loads with no console errors and nothing fetched from outsideNo answer
    Floor Page test
  • Missed.
    Every word and number in the picture is on the pageNo answer
    Floor Page test
  • Missed.
    Wide screen: full-height left sidebar, three cards in a row, table cells in the right columnsNo answer
    Middle Page test
  • Missed.
    Sidebar, active item, accent blue and the three status pills match the pictureNo answer
    Middle Page test
  • Missed.
    Chart bars run Mon to Sun with heights in proportion to their valuesNo answer
    Top Page test
  • Missed.
    On a phone: no sideways scroll, cards stacked, sidebar hidden until a menu button opens itNo answer
    Top Page test
  • Missed.
    Side by side with the picture, a designer would accept it as a faithful buildNo answer
    Top Read
  • Took 5 min 28 s.
  • Wrote 32,000 tokens, 32,000 of them thinking.
  • 65¢ at API prices.
All numbers.
Time to finish5 min 28 s
Time to first wordsnot recorded
Input tokens2
Cached tokens2,327
Output tokens32,000
Thinking tokens32,000
Words in the answer0
Price at API rates65¢
B03A kanban board that survives a reloadLow 7/9Max 0/9Max −7

What it asks for. Asks for a personal kanban board in one HTML file with three columns: to do, doing and done. Cards are added by typing a title, can be dragged between and within columns, edited and deleted, and each column shows how many cards it holds. The board must survive a reload, work from the keyboard with the arrow keys, show titles that contain code exactly as typed, and fit a phone screen.

Max effort passed 7 fewer checks than low.

AI answer from Opus 5.5 · low effort · 2 Oct 2026

7 of 9

Screenshot of the page built in this answer, at a phone width of 375 pixels
Phone, 375 by 812.
Show the code and the answer.

Here's the complete file. Save it as something like kanban.html and open it in a browser.

<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Kanban</title>
<style>
  *, *::before, *::after { box-sizing: border-box; }
  html, body { margin: 0; overflow-x: hidden; }
  body {
    font-family: system-ui, -apple-system, "Segoe UI", Roboto, sans-serif;
    background: #eef1f5;
    color: #1d2330;
    padding: 12px;
  }
  h1 { font-size: 1.3rem; margin: 0 0 12px; }

  .add {
    display: flex;
    gap: 8px;
    margin-bottom: 14px;
    max-width: 600px;
  }
  .add input { flex: 1 1 auto; min-width: 0; }

  input[type="text"] {
    font: inherit;
    padding: 8px 10px;
    border: 1px solid #b7bfcc;
    border-radius: 6px;
    background: #fff;
    width: 100%;
    min-width: 0;
  }
  button {
    font: inherit;
    padding: 7px 12px;
    border: 1px solid #8a94a6;
    border-radius: 6px;
    background: #fff;
    cursor: pointer;
    flex: 0 0 auto;
  }
  button:hover { background: #f1f4f8; }
  button.primary { background: #2f6fde; border-color: #2f6fde; color: #fff; }
  button.primary:hover { background: #255fc4; }

  .board {
    display: grid;
    grid-template-columns: repeat(3, minmax(0, 1fr));
    gap: 12px;
    align-items: start;
  }
  @media (max-width: 700px) {
    .board { grid-template-columns: minmax(0, 1fr); }
  }

  .column {
    background: #dde3eb;
    border-radius: 10px;
    padding: 10px;
    min-width: 0;
  }
  .column h2 {
    font-size: 1rem;
    margin: 0 0 8px;
    display: flex;
    gap: 6px;
    align-items: baseline;
  }
  .count {
    font-size: .85rem;
    background: #fff;
    border-radius: 10px;
    padding: 0 8px;
    color: #445;
  }
  .list {
    list-style: none;
    margin: 0;
    padding: 0;
    min-height: 48px;
    display: flex;
    flex-direction: column;
    gap: 8px;
  }

  .card {
    background: #fff;
    border-radius: 8px;
    padding: 8px;
    box-shadow: 0 1px 2px rgba(0,0,0,.15);
    display: flex;
    gap: 8px;
    align-items: flex-start;
    min-width: 0;
    cursor: grab;
    user-select: none;
    -webkit-user-select: none;
  }
  .card:focus { outline: 3px solid #2f6fde; outline-offset: 2px; }
  .card.dragging { opacity: .6; box-shadow: 0 4px 12px rgba(0,0,0,.25); cursor: grabbing; }

  .handle {
    flex: 0 0 auto;
    color: #8a94a6;
    padding: 2px 4px;
    touch-action: none;   /* lets touch dragging start from the handle */
    cursor: grab;
    line-height: 1.4;
  }
  .body { flex: 1 1 auto; min-width: 0; }
  .title {
    margin: 2px 0 8px;
    white-space: pre-wrap;
    overflow-wrap: anywhere;
    word-break: break-word;
  }
  .actions { display: flex; gap: 6px; flex-wrap: wrap; }
  .actions button { padding: 4px 10px; font-size: .9rem; }
  .card input[type="text"] { margin-bottom: 8px; }

  .empty { color: #6a7385; font-size: .9rem; padding: 6px 2px; }
  .hint { color: #5a6375; font-size: .85rem; margin: 14px 0 0; }
</style>
</head>
<body>
<h1>My tasks</h1>

<form class="add" id="addForm" autocomplete="off">
  <label for="newTitle" style="position:absolute;left:-9999px">New card title</label>
  <input type="text" id="newTitle" placeholder="New card title">
  <button type="submit" class="primary">Add card</button>
</form>

<div class="board" id="board"></div>

<p class="hint">Tip: Tab to a card and use the arrow keys to move it. On touch screens, drag cards by the ⠿ handle.</p>

<script>
(function () {
  "use strict";

  var STORAGE_KEY = "simple-kanban-v1";
  var COLUMNS = [
    { key: "todo",  name: "To do" },
    { key: "doing", name: "Doing" },
    { key: "done",  name: "Done" }
  ];

  var state = load();
  var editingId = null;
  var drag = null;

  var board = document.getElementById("board");
  var addForm = document.getElementById("addForm");
  var newTitle = document.getElementById("newTitle");

  // ---------- storage ----------
  function emptyState() { return { todo: [], doing: [], done: [] }; }

  function load() {
    var s = emptyState();
    try {
      var raw = localStorage.getItem(STORAGE_KEY);
      if (!raw) return s;
      var data = JSON.parse(raw);
      COLUMNS.forEach(function (c) {
        if (Array.isArray(data[c.key])) {
          s[c.key] = data[c.key].filter(function (card) {
            return card && typeof card.id === "string" && typeof card.title === "string";
          }).map(function (card) { return { id: card.id, title: card.title }; });
        }
      });
    } catch (e) { /* bad or missing data: start empty */ }
    return s;
  }

  function save() {
    try { localStorage.setItem(STORAGE_KEY, JSON.stringify(state)); } catch (e) {}
  }

  function newId() {
    return Date.now().toString(36) + Math.random().toString(36).slice(2, 8);
  }

  function find(id) {
    for (var i = 0; i < COLUMNS.length; i++) {
      var list = state[COLUMNS[i].key];
      for (var j = 0; j < list.length; j++) {
        if (list[j].id === id) return { col: i, index: j, card: list[j] };
      }
    }
    return null;
  }

  // ---------- rendering (text only, never HTML) ----------
  function el(tag, className, text) {
    var e = document.createElement(tag);
    if (className) e.className = className;
    if (text != null) e.textContent = text;
    return e;
  }

  function render(focusId, focusEdit) {
    board.textContent = "";
    COLUMNS.forEach(function (col) {
      var cards = state[col.key];
      var section = el("section", "column");
      section.setAttribute("aria-label", col.name);

      var h = el("h2");
      h.appendChild(document.createTextNode(col.name + " "));
      h.appendChild(el("span", "count", String(cards.length)));
      section.appendChild(h);

      var ul = el("ul", "list");
      ul.dataset.col = col.key;
      cards.forEach(function (card) { ul.appendChild(renderCard(card, col)); });
      section.appendChild(ul);
      board.appendChild(section);
    });

    if (focusId) {
      var node = board.querySelector('.card[data-id="' + focusId + '"]');
      if (node) {
        if (focusEdit) {
          var inp = node.querySelector("input");
          if (inp) { inp.focus(); inp.setSelectionRange(inp.value.length, inp.value.length); }
        } else {
          node.focus();
        }
      }
    }
  }

  function renderCard(card, col) {
    var li = el("li", "card");
    li.dataset.id = card.id;
    li.tabIndex = 0;
    li.setAttribute("aria-label", card.title + ", in " + col.name + ". Use arrow keys to move.");

    var handle = el("span", "handle", "⠿");
    handle.setAttribute("aria-hidden", "true");
    li.appendChild(handle);

    var body = el("div", "body");

    if (editingId === card.id) {
      li.setAttribute("aria-label", "Editing " + card.title);
      var input = el("input");
      input.type = "text";
      input.value = card.title;
      input.setAttribute("aria-label", "Card title");
      input.addEventListener("keydown", function (e) {
        if (e.key === "Enter") { e.preventDefault(); commitEdit(card.id, input.value); }
        else if (e.key === "Escape") { e.preventDefault(); editingId = null; render(card.id); }
      });
      body.appendChild(input);

      var a1 = el("div", "actions");
      var saveBtn = el("button", "primary", "Save");
      saveBtn.type = "button";
      saveBtn.addEventListener("click", function () { commitEdit(card.id, input.value); });
      var cancelBtn = el("button", null, "Cancel");
      cancelBtn.type = "button";
      cancelBtn.addEventListener("click", function () { editingId = null; render(card.id); });
      a1.appendChild(saveBtn);
      a1.appendChild(cancelBtn);
      body.appendChild(a1);
    } else {
      body.appendChild(el("p", "title", card.title));

      var a2 = el("div", "actions");
      var editBtn = el("button", null, "Edit");
      editBtn.type = "button";
      editBtn.addEventListener("click", function () {
        editingId = card.id;
        render(card.id, true);
      });
      var delBtn = el("button", null, "Delete");
      delBtn.type = "button";
      delBtn.addEventListener("click", function () { deleteCard(card.id); });
      a2.appendChild(editBtn);
      a2.appendChild(delBtn);
      body.appendChild(a2);
    }
    li.appendChild(body);

    li.addEventListener("keydown", onCardKey);
    li.addEventListener("pointerdown", onPointerDown);
    return li;
  }

  // ---------- actions ----------
  addForm.addEventListener("submit", function (e) {
    e.preventDefault();
    var title = newTitle.value;
    if (title.trim() === "") { newTitle.value = ""; return; }
    state.todo.push({ id: newId(), title: title.trim() });
    save();
    newTitle.value = "";
    render();
    newTitle.focus();
  });

  function commitEdit(id, value) {
    var f = find(id);
    if (f && value.trim() !== "") f.card.title = value.trim();
    editingId = null;
    save();
    render(id);
  }

  function deleteCard(id) {
    var f = find(id);
    if (!f) return;
    var list = state[COLUMNS[f.col].key];
    list.splice(f.index, 1);
    if (editingId === id) editingId = null;
    save();
    // keep keyboard users somewhere sensible
    var next = list[f.index] || list[f.index - 1];
    render(next ? next.id : null);
    if (!next) newTitle.focus();
  }

  // ---------- keyboard moving ----------
  function onCardKey(e) {
    if (e.target !== e.currentTarget) return; // ignore keys typed in inputs/buttons
    var id = e.currentTarget.dataset.id;
    var f = find(id);
    if (!f) return;
    var list = state[COLUMNS[f.col].key];

    if (e.key === "ArrowUp" || e.key === "ArrowDown") {
      e.preventDefault();
      var to = f.index + (e.key === "ArrowUp" ? -1 : 1);
      if (to < 0 || to >= list.length) return;
      list.splice(f.index, 1);
      list.splice(to, 0, f.card);
    } else if (e.key === "ArrowLeft" || e.key === "ArrowRight") {
      e.preventDefault();
      var c = f.col + (e.key === "ArrowLeft" ? -1 : 1);
      if (c < 0 || c >= COLUMNS.length) return;
      list.splice(f.index, 1);
      var target = state[COLUMNS[c].key];
      target.splice(Math.min(f.index, target.length), 0, f.card);
    } else {
      return;
    }
    save();
    render(id);
  }

  // ---------- pointer dragging (mouse, touch, pen) ----------
  function onPointerDown(e) {
    var card = e.currentTarget;
    if (e.button !== 0) return;
    if (e.target.closest("button, input")) return;
    if (e.pointerType !== "mouse" && !e.target.closest(".handle")) return; // touch: handle only, so page can scroll
    drag = { el: card, id: card.dataset.id, x: e.clientX, y: e.clientY, started: false, pointerId: e.pointerId };
    try { card.setPointerCapture(e.pointerId); } catch (err) {}
    card.addEventListener("pointermove", onPointerMove);
    card.addEventListener("pointerup", onPointerEnd);
    card.addEventListener("pointercancel", onPointerEnd);
  }

  function onPointerMove(e) {
    if (!drag || e.pointerId !== drag.pointerId) return;
    if (!drag.started) {
      if (Math.abs(e.clientX - drag.x) + Math.abs(e.clientY - drag.y) < 6) return;
      drag.started = true;
      drag.el.classList.add("dragging");
    }
    e.preventDefault();

    var under = document.elementFromPoint(e.clientX, e.clientY);
    var column = under && under.closest(".column");
    if (!column) return;
    var list = column.querySelector(".list");

    var before = null;
    var cards = list.querySelectorAll(".card");
    for (var i = 0; i < cards.length; i++) {
      if (cards[i] === drag.el) continue;
      var r = cards[i].getBoundingClientRect();
      if (e.clientY < r.top + r.height / 2) { before = cards[i]; break; }
    }
    if (before) {
      if (before.previousSibling !== drag.el) list.insertBefore(drag.el, before);
    } else if (list.lastElementChild !== drag.el) {
      list.appendChild(drag.el);
    }
  }

  function onPointerEnd(e) {
    if (!drag) return;
    var card = drag.el;
    card.removeEventListener("pointermove", onPointerMove);
    card.removeEventListener("pointerup", onPointerEnd);
    card.removeEventListener("pointercancel", onPointerEnd);
    try { card.releasePointerCapture(drag.pointerId); } catch (err) {}
    var started = drag.started, id = drag.id;
    drag = null;
    if (!started) return;

    // Read the new order back from the page into the saved state
    var all = {};
    COLUMNS.forEach(function (c) { state[c.key].forEach(function (cd) { all[cd.id] = cd; }); });
    var fresh = emptyState();
    board.querySelectorAll(".list").forEach(function (ul) {
      ul.querySelectorAll(".card").forEach(function (li) {
        if (all[li.dataset.id]) fresh[ul.dataset.col].push(all[li.dataset.id]);
      });
    });
    state = fresh;
    save();
    render(id);
  }

  render();
})();
</script>
</body>
</html>

How it covers each point:

  • Columns and counts: Each column header shows its name and a number, for example "Doing 2". The board starts empty.
  • Adding cards: Type a title, then press Enter or click Add card. The card goes to the bottom of To do. Blank or space-only titles are ignored.
  • Dragging: With a mouse, you can drag a card from anywhere on it, both between columns and up and down within one. On a touchscreen, drag by the ⠿ handle. Touching anywhere else on a card still scrolls the page normally.
  • Edit and Delete: Edit turns the title into a text field holding the current title. Enter or Save confirms it. Escape or Cancel backs out. If you save an empty title, the old one is kept. Delete removes the card straight away with no confirmation.
  • Saving: The board is stored in your browser's localStorage after every change, so it looks the same after a reload.
  • Keyboard: Every card is a Tab stop of its own. Left and right arrows move it to the neighbouring column, and up and down move it within its column. Focus stays on the card after each move, so you can keep pressing arrows.
  • Code in titles: Titles are only ever inserted as plain text, never as HTML. Something like <div> or a & b shows exactly as you typed it.
  • Phones: Below 700px wide the columns stack vertically. Long titles, even unbroken strings, wrap inside the card, so the page never scrolls sideways at 375px.
  • Floor2/2, 100% passed
  • Middle3/3, 100% passed
  • Top2/4, 50% passed
7 of 9 checks passedfloor 2/2, middle 3/3, top 2/4
  • Passed.
    The page loads with no errors and has the three named columns
    Floor Page test
  • Passed.
    Adding a card works with Enter and with the Add card button
    Floor Page test
  • Passed.
    Empty or space-only titles are ignored
    Middle Page test
  • Passed.
    The counts update, and Edit and Delete work
    Middle Page test
  • Passed.
    After a reload the board is exactly as it was
    Middle Page test
  • Missed.
    Dragging moves cards between columns and reorders them, and it survives a reloadtest group failed
    Top Page test
  • Missed.
    A card can be moved with the arrow keys and keeps focustest group failed
    Top Page test
  • Passed.
    Titles like <b>hi</b> & <img ...> show exactly as typed and run nothing
    Top Page test
  • Passed.
    No sideways scroll on a 375 px phone, even with very long titles
    Top Page test
  • Took 57 s.
  • First words after 14 s.
  • Wrote 7,037 tokens, 1,086 of them thinking.
  • 15¢ at API prices.
All numbers.
Time to finish57 s
Time to first words14 s
Input tokens2
Cached tokens1,169
Output tokens7,037
Thinking tokens1,086
Words in the answer1,710
Price at API rates15¢

AI answer from Opus 5.5 · max effort · 2 Oct 2026

0 of 9

The answer was cut off at the token limit.

  • Floor0/2, 0% passed
  • Middle0/3, 0% passed
  • Top0/4, 0% passed
0 of 9 checks passedfloor 0/2, middle 0/3, top 0/4
  • Missed.
    The page loads with no errors and has the three named columnsNo answer
    Floor Page test
  • Missed.
    Adding a card works with Enter and with the Add card buttonNo answer
    Floor Page test
  • Missed.
    Empty or space-only titles are ignoredNo answer
    Middle Page test
  • Missed.
    The counts update, and Edit and Delete workNo answer
    Middle Page test
  • Missed.
    After a reload the board is exactly as it wasNo answer
    Middle Page test
  • Missed.
    Dragging moves cards between columns and reorders them, and it survives a reloadNo answer
    Top Page test
  • Missed.
    A card can be moved with the arrow keys and keeps focusNo answer
    Top Page test
  • Missed.
    Titles like <b>hi</b> & <img ...> show exactly as typed and run nothingNo answer
    Top Page test
  • Missed.
    No sideways scroll on a 375 px phone, even with very long titlesNo answer
    Top Page test
  • Took 5 min 27 s.
  • Wrote 32,000 tokens, 31,999 of them thinking.
  • 65¢ at API prices.
All numbers.
Time to finish5 min 27 s
Time to first wordsnot recorded
Input tokens2
Cached tokens1,169
Output tokens32,000
Thinking tokens31,999
Words in the answer0
Price at API rates65¢
B04Minesweeper in one HTML fileLow 8/9Max 0/9Max −8

What it asks for. Asks for the game Minesweeper in one HTML file: a 9 by 9 board with 10 mines, flags on right click, a mines-left counter, a first square that is never a mine, empty areas that open by themselves, win and lose messages and a new game button. Every square has to be a button so the game works from the keyboard, and the board has to fit a phone.

Max effort passed 8 fewer checks than low.

AI answer from Opus 5.5 · low effort · 2 Oct 2026

8 of 9

Screenshot of the page built in this answer, at a desktop width of 1280 pixels
Our screenshot, 1280 by 800. Open it for the full length.
Screenshot of the page built in this answer, at a phone width of 375 pixels
Phone, 375 by 812.
More screenshots (2).
Screenshot of the page built in this answer, at a desktop width of 1280 pixelsScreenshot of the page built in this answer, at a phone width of 375 pixels
Show the code and the answer.

Here's the complete file. Save it as something like minesweeper.html and open it in any browser.

<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Minesweeper</title>
<style>
  * { box-sizing: border-box; }
  body {
    margin: 0;
    padding: 16px 8px;
    font-family: system-ui, -apple-system, "Segoe UI", Roboto, sans-serif;
    background: #eef1f5;
    color: #222;
    display: flex;
    justify-content: center;
  }
  .game {
    width: 100%;
    max-width: 360px;
    text-align: center;
  }
  h1 { font-size: 1.5rem; margin: 0 0 12px; }
  .bar {
    display: flex;
    justify-content: space-between;
    align-items: center;
    margin-bottom: 10px;
    gap: 8px;
  }
  .counter { font-size: 1.05rem; font-weight: 600; }
  #newGame {
    font-size: 1rem;
    padding: 8px 14px;
    min-height: 40px;
    border: 1px solid #888;
    border-radius: 6px;
    background: #fff;
    cursor: pointer;
  }
  #newGame:focus-visible, .cell:focus-visible {
    outline: 3px solid #1a73e8;
    outline-offset: 1px;
  }
  .board {
    display: grid;
    grid-template-columns: repeat(9, 36px);
    gap: 2px;
    justify-content: center;
    margin: 0 auto;
    user-select: none;
    -webkit-user-select: none;
  }
  .cell {
    width: 36px;
    height: 36px;
    padding: 0;
    font-size: 1.1rem;
    font-weight: 700;
    border: 1px solid #8a94a3;
    border-radius: 3px;
    background: #c3cad4;
    cursor: pointer;
    line-height: 1;
    font-family: inherit;
    touch-action: manipulation;
  }
  .cell:hover:not(:disabled):not(.open) { background: #b3bbc7; }
  .cell.open {
    background: #fff;
    border-color: #d0d5dc;
    cursor: default;
  }
  .cell.boom { background: #f28b82; }
  .cell:disabled { cursor: default; color: inherit; opacity: 1; }
  .n1 { color: #1a56db; } .n2 { color: #127a2e; } .n3 { color: #c81e1e; }
  .n4 { color: #1e2a78; } .n5 { color: #7a1e1e; } .n6 { color: #0f7c7c; }
  .n7 { color: #000; }    .n8 { color: #666; }
  #message {
    min-height: 1.6em;
    margin-top: 12px;
    font-size: 1.3rem;
    font-weight: 700;
  }
  .help { font-size: 0.85rem; color: #555; margin-top: 8px; }
</style>
</head>
<body>
<main class="game">
  <h1>Minesweeper</h1>
  <div class="bar">
    <div class="counter">Mines left: <span id="minesLeft">10</span></div>
    <button id="newGame" type="button">New game</button>
  </div>
  <div id="board" class="board" role="grid" aria-label="Minesweeper board"></div>
  <div id="message" role="status" aria-live="polite"></div>
  <p class="help">Click or press Enter/Space to open. Right click (or press F) to place or remove a flag.</p>
</main>

<script>
(function () {
  const ROWS = 9, COLS = 9, MINES = 10;
  const boardEl = document.getElementById('board');
  const counterEl = document.getElementById('minesLeft');
  const messageEl = document.getElementById('message');
  const newGameBtn = document.getElementById('newGame');

  let cells = [];       // {mine, open, flag, count, el}
  let minesPlaced = false;
  let gameOver = false;
  let flags = 0;
  let openedCount = 0;

  function idx(r, c) { return r * COLS + c; }

  function neighbours(i) {
    const r = Math.floor(i / COLS), c = i % COLS, out = [];
    for (let dr = -1; dr <= 1; dr++) {
      for (let dc = -1; dc <= 1; dc++) {
        if (dr === 0 && dc === 0) continue;
        const nr = r + dr, nc = c + dc;
        if (nr >= 0 && nr < ROWS && nc >= 0 && nc < COLS) out.push(idx(nr, nc));
      }
    }
    return out;
  }

  function label(i) {
    const r = Math.floor(i / COLS) + 1, c = (i % COLS) + 1;
    const cell = cells[i];
    let state = 'unopened';
    if (cell.flag) state = 'flagged';
    else if (cell.open) state = cell.mine ? 'mine' : (cell.count ? cell.count + ' adjacent mines' : 'empty');
    else if (gameOver && cell.mine) state = 'mine';
    cell.el.setAttribute('aria-label', 'Row ' + r + ', column ' + c + ', ' + state);
  }

  function newGame() {
    boardEl.innerHTML = '';
    cells = [];
    minesPlaced = false;
    gameOver = false;
    flags = 0;
    openedCount = 0;
    messageEl.textContent = '';
    updateCounter();

    for (let i = 0; i < ROWS * COLS; i++) {
      const btn = document.createElement('button');
      btn.type = 'button';
      btn.className = 'cell';
      btn.dataset.i = i;
      const cell = { mine: false, open: false, flag: false, count: 0, el: btn };
      cells.push(cell);
      boardEl.appendChild(btn);
      label(i);
    }
  }

  function placeMines(safe) {
    const spots = [];
    for (let i = 0; i < ROWS * COLS; i++) if (i !== safe) spots.push(i);
    // Fisher–Yates shuffle, take the first MINES
    for (let i = spots.length - 1; i > 0; i--) {
      const j = Math.floor(Math.random() * (i + 1));
      [spots[i], spots[j]] = [spots[j], spots[i]];
    }
    for (let k = 0; k < MINES; k++) cells[spots[k]].mine = true;
    for (let i = 0; i < cells.length; i++) {
      cells[i].count = neighbours(i).filter(n => cells[n].mine).length;
    }
    minesPlaced = true;
  }

  function updateCounter() {
    counterEl.textContent = MINES - flags;
  }

  function reveal(i) {
    const cell = cells[i];
    cell.open = true;
    openedCount++;
    cell.el.classList.add('open');
    if (cell.count > 0) {
      cell.el.textContent = cell.count;
      cell.el.classList.add('n' + cell.count);
    } else {
      cell.el.textContent = '';
    }
    label(i);
  }

  function openCell(i) {
    if (gameOver) return;
    const cell = cells[i];
    if (cell.open || cell.flag) return;

    if (!minesPlaced) placeMines(i);

    if (cell.mine) {
      lose(i);
      return;
    }

    // Iterative flood fill
    const stack = [i];
    while (stack.length) {
      const cur = stack.pop();
      const c = cells[cur];
      if (c.open || c.flag || c.mine) continue;
      reveal(cur);
      if (c.count === 0) {
        for (const n of neighbours(cur)) {
          if (!cells[n].open && !cells[n].flag) stack.push(n);
        }
      }
    }

    if (openedCount === ROWS * COLS - MINES) win();
  }

  function toggleFlag(i) {
    if (gameOver) return;
    const cell = cells[i];
    if (cell.open) return;
    cell.flag = !cell.flag;
    flags += cell.flag ? 1 : -1;
    cell.el.textContent = cell.flag ? '🚩' : '';
    updateCounter();
    label(i);
  }

  function endBoard() {
    gameOver = true;
    for (const c of cells) c.el.disabled = true;
  }

  function lose(hit) {
    gameOver = true;
    for (let i = 0; i < cells.length; i++) {
      const c = cells[i];
      if (c.mine) {
        c.el.textContent = '💣';
        c.el.classList.add('open');
      }
      label(i);
    }
    cells[hit].el.classList.add('boom');
    messageEl.textContent = 'Game over';
    endBoard();
  }

  function win() {
    messageEl.textContent = 'You win!';
    endBoard();
  }

  // Left click, and Enter/Space (buttons fire "click" for those keys)
  boardEl.addEventListener('click', function (e) {
    const btn = e.target.closest('.cell');
    if (!btn) return;
    openCell(Number(btn.dataset.i));
  });

  // Right click toggles a flag
  boardEl.addEventListener('contextmenu', function (e) {
    e.preventDefault();
    const btn = e.target.closest('.cell');
    if (!btn) return;
    toggleFlag(Number(btn.dataset.i));
  });

  // Keyboard extras: F to flag, arrow keys to move between squares
  boardEl.addEventListener('keydown', function (e) {
    const btn = e.target.closest('.cell');
    if (!btn) return;
    const i = Number(btn.dataset.i);
    const r = Math.floor(i / COLS), c = i % COLS;
    let target = null;
    switch (e.key) {
      case 'f': case 'F': e.preventDefault(); toggleFlag(i); return;
      case 'ArrowUp':    if (r > 0) target = idx(r - 1, c); break;
      case 'ArrowDown':  if (r < ROWS - 1) target = idx(r + 1, c); break;
      case 'ArrowLeft':  if (c > 0) target = idx(r, c - 1); break;
      case 'ArrowRight': if (c < COLS - 1) target = idx(r, c + 1); break;
      default: return;
    }
    e.preventDefault();
    if (target !== null) cells[target].el.focus();
  });

  newGameBtn.addEventListener('click', newGame);

  newGame();
})();
</script>
</body>
</html>

How it meets each requirement:

  • Board: The board is 9×9 with 10 mines. Each square is a real <button>, so Tab, Enter and Space work without any extra code.
  • Opening and flagging: Left click opens a square, and Enter or Space opens the focused one. Right click places or removes a 🚩. A flagged square ignores clicks until you take the flag off.
  • Extra keyboard controls: Right click has no keyboard equivalent, so I added two things. Pressing F toggles a flag on the focused square. The arrow keys move between squares.
  • Mines left: The counter shows 10 − flags.
  • First click is safe: Mines are only placed after your first open, and that square is left out.
  • Numbers and auto-opening: Each number is the count of the 8 neighbours that hold a mine. Opening a square with no mines around it opens its neighbours, and this keeps spreading. It uses a stack rather than recursion.
  • Winning: When every non-mine square is open, the game shows "You win!" and the board locks.
  • Losing: Opening a mine shows "Game over" and reveals every mine as 💣. The mine you hit is highlighted in red, and the board's buttons are disabled.
  • New game: This button makes a fresh board with new mines, sets the counter back to 10 and clears the message.
  • Size: Each square is 36px. The board is 9 × 36px plus 8 × 2px gaps, which is 340px wide. It fits inside a 375px screen without sideways scrolling.
  • Floor2/2, 100% passed
  • Middle2/3, 67% passed
  • Top4/4, 100% passed
8 of 9 checks passedfloor 2/2, middle 2/3, top 4/4
  • Passed.
    81 square buttons in a 9 by 9 grid, Mines left starts at 10
    Floor Page test
  • Passed.
    A first click never lands on a mine
    Floor Page test
  • Missed.
    Numbers are right and an empty square opens its neighbourstest group failed
    Middle Page test
  • Passed.
    Right click flags and unflags, the counter follows, a flagged square stays shut
    Middle Page test
  • Passed.
    The squares work from the keyboard
    Middle Page test
  • Passed.
    Hitting a mine says Game over, shows all 10 mines as 💣 and stops the board
    Top Page test
  • Passed.
    Opening every safe square shows You win! (and not before)
    Top Page test
  • Passed.
    New game resets the board, counter and message
    Top Page test
  • Passed.
    Fits a 375 px phone with squares of at least 32 px, no sideways scroll
    Top Page test
  • Took 37 s.
  • First words after 6.6 s.
  • Wrote 4,595 tokens, 381 of them thinking.
  • 9.7¢ at API prices.
All numbers.
Time to finish37 s
Time to first words6.6 s
Input tokens2
Cached tokens1,111
Output tokens4,595
Thinking tokens381
Words in the answer1,332
Price at API rates9.7¢

AI answer from Opus 5.5 · max effort · 2 Oct 2026

0 of 9

The answer was cut off at the token limit.

  • Floor0/2, 0% passed
  • Middle0/3, 0% passed
  • Top0/4, 0% passed
0 of 9 checks passedfloor 0/2, middle 0/3, top 0/4
  • Missed.
    81 square buttons in a 9 by 9 grid, Mines left starts at 10No answer
    Floor Page test
  • Missed.
    A first click never lands on a mineNo answer
    Floor Page test
  • Missed.
    Numbers are right and an empty square opens its neighboursNo answer
    Middle Page test
  • Missed.
    Right click flags and unflags, the counter follows, a flagged square stays shutNo answer
    Middle Page test
  • Missed.
    The squares work from the keyboardNo answer
    Middle Page test
  • Missed.
    Hitting a mine says Game over, shows all 10 mines as 💣 and stops the boardNo answer
    Top Page test
  • Missed.
    Opening every safe square shows You win! (and not before)No answer
    Top Page test
  • Missed.
    New game resets the board, counter and messageNo answer
    Top Page test
  • Missed.
    Fits a 375 px phone with squares of at least 32 px, no sideways scrollNo answer
    Top Page test
  • Took 5 min 6 s.
  • Wrote 32,000 tokens, 32,000 of them thinking.
  • 65¢ at API prices.
All numbers.
Time to finish5 min 6 s
Time to first wordsnot recorded
Input tokens2
Cached tokens1,111
Output tokens32,000
Thinking tokens32,000
Words in the answer0
Price at API rates65¢
B05Fix three layout bugs on a garden pageLow 7/7Max 0/7Max −7

What it asks for. Pastes a community garden club's web page and describes three problems on phones: the page scrolls sideways, section headings end up hidden under the sticky header after tapping the menu links, and three event cards are squashed instead of sitting one per row. It asks for the whole page back with all three fixed, leaving the laptop layout as it is and the header still sticky.

Max effort passed 7 fewer checks than low.

AI answer from Opus 5.5 · low effort · 2 Oct 2026

7 of 7

Screenshot of the page built in this answer, at a desktop width of 1280 pixels
Our screenshot, 1280 by 800. Open it for the full length.
Screenshot of the page built in this answer, at a phone width of 375 pixels
Phone, 375 by 812.
Show the code and the answer.

Here's what was going wrong on phones, and the fixed file.

1. Sideways scrolling. Three things were wider than a phone screen:

  • The fees table had a fixed width: 720px. It now fills the available width up to 720px, so laptops look the same. It also sits in a wrapper that scrolls on its own if a very narrow screen still can't fit it, so only the table scrolls, not the whole page.
  • The long email address is one unbreakable "word". It's now allowed to wrap.
  • The nav links couldn't wrap on very small screens. Now they can.

2. Headings hidden under the sticky header. The browser scrolls the section to the very top of the screen, which sits behind the header. scroll-margin-top tells it to stop short by the header's height. On phones the header wraps onto two lines, so it gets a larger value there.

3. Squashed cards. On screens 640px wide or less, the grid switches to one card per row. I also changed 1fr to minmax(0, 1fr) so long words can't force the cards wider than their columns.

Laptop layout and the sticky header are unchanged.

<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Thistlebank Allotments</title>
<style>
  * { box-sizing: border-box; }
  body {
    margin: 0;
    font-family: system-ui, -apple-system, "Segoe UI", Helvetica, Arial, sans-serif;
    line-height: 1.55;
    color: #23301f;
    background: #f7f4ea;
  }
  .wrap { max-width: 960px; margin: 0 auto; padding: 0 1rem; }
  .site-header {
    position: sticky;
    top: 0;
    z-index: 10;
    background: #2f5d34;
    color: #fff;
    box-shadow: 0 1px 6px rgba(0, 0, 0, .25);
  }
  .site-header .wrap {
    display: flex;
    flex-wrap: wrap;
    align-items: center;
    justify-content: space-between;
    gap: .25rem 1.5rem;
    padding-top: .9rem;
    padding-bottom: .9rem;
  }
  .brand { font-size: 1.15rem; font-weight: 700; letter-spacing: .01em; }
  .site-header nav { display: flex; flex-wrap: wrap; gap: .25rem 1.1rem; }
  .site-header a { color: #fff; text-decoration: none; font-weight: 500; }
  .site-header a:hover { text-decoration: underline; }
  .hero { background: #dfe9c8; padding: 4rem 0 3.5rem; }
  .hero h1 { margin: 0 0 .5rem; font-size: 2.4rem; line-height: 1.15; color: #1f4424; }
  .hero p { margin: 0; max-width: 34rem; font-size: 1.15rem; }
  section.block { padding: 2rem 0; scroll-margin-top: 4rem; }
  section.block h2 { margin: 0 0 1rem; font-size: 1.7rem; color: #1f4424; }
  .cards {
    display: grid;
    grid-template-columns: repeat(3, minmax(0, 1fr));
    gap: 1.25rem;
  }
  .card { background: #fff; border: 1px solid #d8d2bd; border-radius: 10px; padding: 1.1rem 1.2rem; }
  .card h3 { margin: 0 0 .2rem; font-size: 1.15rem; }
  .card .when { margin: 0 0 .6rem; font-size: .9rem; color: #5b6650; font-weight: 600; }
  .card p { margin: 0; }
  .table-scroll { max-width: 100%; overflow-x: auto; }
  .fees { width: 100%; max-width: 720px; border-collapse: collapse; background: #fff; }
  .fees th, .fees td { text-align: left; padding: .7rem .9rem; border-bottom: 1px solid #e2dcc7; }
  .fees th { background: #e8e2cc; font-size: .95rem; }
  .note { margin-top: 1rem; font-size: .95rem; color: #4a5642; }
  .contact-box { background: #fff; border: 1px solid #d8d2bd; border-radius: 10px; padding: 1.2rem 1.4rem; }
  .contact-box p { margin: 0 0 .5rem; }
  .email { font-size: 1.1rem; font-weight: 600; overflow-wrap: anywhere; }
  .email a { color: #1f4424; }
  footer { background: #23301f; color: #d8dccc; padding: 1.5rem 0; font-size: .9rem; }

  @media (max-width: 640px) {
    .cards { grid-template-columns: 1fr; }
    /* The header wraps onto two lines here, so it is taller */
    section.block { scroll-margin-top: 6.5rem; }
    .fees th, .fees td { padding: .6rem .6rem; }
  }
</style>
</head>
<body>
<header class="site-header">
  <div class="wrap">
    <span class="brand">Thistlebank Allotments</span>
    <nav>
      <a href="#about">About</a>
      <a href="#events">Events</a>
      <a href="#fees">Plot fees</a>
      <a href="#contact">Contact</a>
    </nav>
  </div>
</header>

<main>
  <div class="hero">
    <div class="wrap">
      <h1>Grow something together</h1>
      <p>A friendly community garden on the edge of town, with forty-two plots, a shared tool shed and a kettle that is always on.</p>
    </div>
  </div>

  <section class="block" id="about">
    <div class="wrap">
      <h2>About the garden</h2>
      <p>Thistlebank has been run by its members since the garden opened. Plots are small enough to manage in an evening a week, and there is a shared bed for anyone who is not ready to commit to a plot of their own.</p>
      <p>Beginners are welcome. Long-time growers are usually happy to explain why your beans are doing that.</p>
    </div>
  </section>

  <section class="block" id="events">
    <div class="wrap">
      <h2>Coming up</h2>
      <div class="cards">
        <article class="card">
          <h3>Seed swap</h3>
          <p class="when">First Saturday of the month, 10:00</p>
          <p>Bring spare seeds, take what you need. Tea and biscuits in the shed.</p>
        </article>
        <article class="card">
          <h3>Compost clinic</h3>
          <p class="when">Second Wednesday, 18:30</p>
          <p>Learn why the heap smells and what to do about it. Bring gloves.</p>
        </article>
        <article class="card">
          <h3>Autumn bake and dig</h3>
          <p class="when">Last Sunday of October, 13:00</p>
          <p>Tidy the shared beds, then eat cake made from whatever we grew.</p>
        </article>
      </div>
    </div>
  </section>

  <section class="block" id="fees">
    <div class="wrap">
      <h2>Plot fees</h2>
      <div class="table-scroll">
        <table class="fees">
          <thead>
            <tr><th>Plot size</th><th>Best for</th><th>Per year</th><th>Water</th></tr>
          </thead>
          <tbody>
            <tr><td>Small (3 x 5 m)</td><td>Salad and herbs</td><td>$35</td><td>Included</td></tr>
            <tr><td>Medium (5 x 8 m)</td><td>A family of four</td><td>$60</td><td>Included</td></tr>
            <tr><td>Large (8 x 10 m)</td><td>Serious growers</td><td>$95</td><td>Included</td></tr>
            <tr><td>Shared bed</td><td>Trying it out</td><td>$10</td><td>Included</td></tr>
          </tbody>
        </table>
      </div>
      <p class="note">Fees are due on the first of March. Concessions are available; just ask.</p>
    </div>
  </section>

  <section class="block" id="contact">
    <div class="wrap">
      <h2>Contact</h2>
      <div class="contact-box">
        <p>Write to the plot secretary about the waiting list, a visit or anything else.</p>
        <p class="email"><a href="mailto:membershipandplotenquiries@thistlebankallotments.example">membershipandplotenquiries@thistlebankallotments.example</a></p>
        <p>Open days are the first Saturday of every month.</p>
      </div>
    </div>
  </section>
</main>

<footer>
  <div class="wrap">Thistlebank Allotments. Run by members, for members.</div>
</footer>
</body>
</html>

If you later change the header (bigger text, more links), adjust the two scroll-margin-top values (4rem and 6.5rem) to match its new height.

  • Floor2/2, 100% passed
  • Middle2/2, 100% passed
  • Top3/3, 100% passed
7 of 7 checks passedfloor 2/2, middle 2/2, top 3/3
  • Passed.
    Returns one whole HTML page that loads
    Floor Page test
  • Passed.
    All the original text is still on the page
    Floor Page test
  • Passed.
    On a 375 px phone the page no longer scrolls sideways
    Middle Page test
  • Passed.
    On a phone the three event cards stack, one per row
    Middle Page test
  • Passed.
    Tapping each top link shows the section heading below the header (phone and laptop)
    Top Page test
  • Passed.
    Laptop layout unchanged: header still sticky, three cards in a row, blocks in the same places
    Top Page test
  • Passed.
    No console errors and nothing loaded from outside
    Top Page test
  • Took 30 s.
  • First words after 9.1 s.
  • Wrote 3,732 tokens, 623 of them thinking.
  • 9.1¢ at API prices.
All numbers.
Time to finish30 s
Time to first words9.1 s
Input tokens2
Cached tokens3,361
Output tokens3,732
Thinking tokens623
Words in the answer907
Price at API rates9.1¢

AI answer from Opus 5.5 · max effort · 2 Oct 2026

0 of 7

The answer was cut off at the token limit.

  • Floor0/2, 0% passed
  • Middle0/2, 0% passed
  • Top0/3, 0% passed
0 of 7 checks passedfloor 0/2, middle 0/2, top 0/3
  • Missed.
    Returns one whole HTML page that loadsNo answer
    Floor Page test
  • Missed.
    All the original text is still on the pageNo answer
    Floor Page test
  • Missed.
    On a 375 px phone the page no longer scrolls sidewaysNo answer
    Middle Page test
  • Missed.
    On a phone the three event cards stack, one per rowNo answer
    Middle Page test
  • Missed.
    Tapping each top link shows the section heading below the header (phone and laptop)No answer
    Top Page test
  • Missed.
    Laptop layout unchanged: header still sticky, three cards in a row, blocks in the same placesNo answer
    Top Page test
  • Missed.
    No console errors and nothing loaded from outsideNo answer
    Top Page test
  • Took 5 min 5 s.
  • Wrote 32,000 tokens, 32,000 of them thinking.
  • 66¢ at API prices.
All numbers.
Time to finish5 min 5 s
Time to first wordsnot recorded
Input tokens2
Cached tokens3,361
Output tokens32,000
Thinking tokens32,000
Words in the answer0
Price at API rates66¢

Code

F01Load more shows some posts twice and skips othersLow 7/7Max 7/7Same

What it asks for. Pastes a Python function that pages through a community board's posts, newest first, and reports that some posts show up twice while others never appear, more often on busy days. Many posts share exactly the same timestamp. It asks for the cause and a fixed function with the same name, arguments and return shape, where posts with the same time come higher id first.

Both settings passed 7 of 7 checks.

AI answer from Opus 5.5 · low effort · 2 Oct 2026

7 of 7

What was wrong

  1. The cursor only stores created_at, and the query uses <=. Every post on the next page that shares the last post's timestamp gets fetched again, so you see duplicates.
  2. The order among equal timestamps is undefined. ORDER BY created_at DESC lets SQLite return tied rows in any order, and that order can differ between queries. Some tied posts fall into the gap between pages and are never shown.
  3. Switching to < would not fix it. It would skip every remaining post with the boundary timestamp.

This happens more on busy days, and with the importer, because both produce many posts with the same timestamp.

The fix

The fix is keyset pagination on the pair (created_at, id):

  • Sort by created_at DESC, id DESC. Since id is unique, this gives a total, stable order.
  • Put both values in the cursor.
  • Ask for rows that come strictly after the cursor position in that order.
import base64
import sqlite3

MAX_LIMIT = 100
COLUMNS = "id, author, body, created_at"


def _encode_cursor(created_at, post_id):
    raw = f"{int(created_at)}:{int(post_id)}".encode("ascii")
    return base64.urlsafe_b64encode(raw).decode("ascii")


def _decode_cursor(cursor):
    try:
        raw = base64.urlsafe_b64decode(cursor.encode("ascii")).decode("ascii")
        created_at, post_id = raw.split(":")
        return int(created_at), int(post_id)
    except (ValueError, UnicodeError, AttributeError) as exc:
        raise ValueError("invalid cursor") from exc


def fetch_page(conn, cursor=None, limit=20):
    """Return (rows, next_cursor): one page of the board, newest first.

    rows is a list of (id, author, body, created_at) tuples, ordered by
    created_at DESC, then id DESC for posts sharing a created_at.
    next_cursor is an opaque string the caller hands back to get the next
    page, or None when there are no more posts.

    posts(id INTEGER PRIMARY KEY, author TEXT, body TEXT, created_at INTEGER)
    created_at is unix seconds.
    """
    limit = max(1, min(int(limit), MAX_LIMIT))

    if cursor is None:
        cur = conn.execute(
            f"SELECT {COLUMNS} FROM posts "
            "ORDER BY created_at DESC, id DESC LIMIT ?",
            (limit + 1,),
        )
    else:
        last_created_at, last_id = _decode_cursor(cursor)
        cur = conn.execute(
            f"SELECT {COLUMNS} FROM posts "
            "WHERE created_at < ? OR (created_at = ? AND id < ?) "
            "ORDER BY created_at DESC, id DESC LIMIT ?",
            (last_created_at, last_created_at, last_id, limit + 1),
        )

    rows = cur.fetchall()
    has_more = len(rows) > limit
    rows = rows[:limit]

    if has_more:
        last = rows[-1]
        next_cursor = _encode_cursor(last[3], last[0])
    else:
        next_cursor = None
    return rows, next_cursor

Notes

  • Index: add CREATE INDEX IF NOT EXISTS idx_posts_created_id ON posts(created_at DESC, id DESC); so each page is a quick index seek rather than a sort of the whole table.
  • Cursor format: the cursor is now base64 of "created_at:id". The app still just passes it back. A malformed cursor raises ValueError instead of producing a silently wrong page.
  • Old cursors: cursors issued by the old version will fail to decode. Clients will need to restart from the first page once after you deploy.
  • Alternative query: on SQLite 3.15 or later you could write the condition as WHERE (created_at, id) < (?, ?). The expanded OR form above does the same thing and works on all versions.
  • Floor2/2, 100% passed
  • Middle3/3, 100% passed
  • Top2/2, 100% passed
7 of 7 checks passedfloor 2/2, middle 3/3, top 2/2
  • Passed.
    The first page comes back correct, newest first
    Floor Tested
  • Passed.
    Keeps the return shape: string cursor, None at the end
    Floor Tested
  • Passed.
    Paging through everything shows each post exactly once
    Middle Tested
  • Passed.
    Order is newest first and, on ties, higher id first
    Middle Tested
  • Passed.
    Explains that shared created_at values cause the problem
    Middle Tested
  • Passed.
    Still correct at page sizes 1, 2 and 3
    Top Tested
  • Passed.
    New posts added at the top mid-scroll cause no repeats or gaps
    Top Tested
  • Took 15 s.
  • First words after 4.8 s.
  • Wrote 1,563 tokens, 165 of them thinking.
  • 3.8¢ at API prices.
All numbers.
Time to finish15 s
Time to first words4.8 s
Input tokens2
Cached tokens1,408
Output tokens1,563
Thinking tokens165
Words in the answer500
Price at API rates3.8¢

AI answer from Opus 5.5 · max effort · 2 Oct 2026

7 of 7

What was wrong

The cursor is only the last post's created_at, and that value isn't unique.

  • Duplicates: WHERE created_at <= cursor fetches the last post of the previous page again, along with every other post that has the same timestamp. Busy days mean more posts in the same second, so you get more repeats.
  • Missing posts: if a full page or more of posts share one timestamp (easy with the importer), the next page contains only those posts. It ends on the same timestamp, so the cursor never moves. The same page keeps coming back, and nothing after it is ever reached.
  • No tie-breaker: with only ORDER BY created_at, the order of tied posts is undefined. SQLite can return them in a different order on each query, so which page they land on is arbitrary.
  • Changing <= to < would only swap the duplicates for skipped posts.

Fix: sort by created_at DESC, id DESC. Because ids are unique, this gives every post a fixed place. The cursor stores both values of the last row sent, and the next page starts strictly after that exact row. Every post then appears exactly once, and a large group of same-second posts can't stall the scroll.

import sqlite3

MAX_LIMIT = 100
COLUMNS = "id, author, body, created_at"


def fetch_page(conn, cursor=None, limit=20):
    """Return (rows, next_cursor): one page of the board, newest first.

    Posts with the same created_at are ordered by id, highest first.

    rows is a list of (id, author, body, created_at) tuples.
    next_cursor is an opaque string the caller hands back to get the next
    page, or None when there are no more posts.

    posts(id INTEGER PRIMARY KEY, author TEXT, body TEXT, created_at INTEGER)
    created_at is unix seconds.
    """
    limit = max(1, min(int(limit), MAX_LIMIT))

    if cursor is None:
        cur = conn.execute(
            f"SELECT {COLUMNS} FROM posts "
            "ORDER BY created_at DESC, id DESC LIMIT ?",
            (limit + 1,),
        )
    else:
        # The cursor is "created_at:id" of the last row already returned.
        # Return rows strictly after it in (created_at DESC, id DESC) order.
        last_created_at, last_id = map(int, cursor.split(":"))
        cur = conn.execute(
            f"SELECT {COLUMNS} FROM posts "
            "WHERE created_at <= ? AND (created_at < ? OR id < ?) "
            "ORDER BY created_at DESC, id DESC LIMIT ?",
            (last_created_at, last_created_at, last_id, limit + 1),
        )

    rows = cur.fetchall()
    has_more = len(rows) > limit
    rows = rows[:limit]

    next_cursor = f"{rows[-1][3]}:{rows[-1][0]}" if has_more else None
    return rows, next_cursor

Two small notes

  • Old cursors: cursors handed out by the old version are a bare number. This version rejects them with ValueError, the same as any malformed cursor. Anyone mid-scroll during the deploy will get one failed "load more" and need to reload. If that matters, you could run the old created_at <= ? query for a cursor with no : for that one request.
  • Performance: an index on posts(created_at) is all this query needs. SQLite stores the rowid (your id) at the end of every index entry, so it can walk that index in order without a separate sort.
  • Floor2/2, 100% passed
  • Middle3/3, 100% passed
  • Top2/2, 100% passed
7 of 7 checks passedfloor 2/2, middle 3/3, top 2/2
  • Passed.
    The first page comes back correct, newest first
    Floor Tested
  • Passed.
    Keeps the return shape: string cursor, None at the end
    Floor Tested
  • Passed.
    Paging through everything shows each post exactly once
    Middle Tested
  • Passed.
    Order is newest first and, on ties, higher id first
    Middle Tested
  • Passed.
    Explains that shared created_at values cause the problem
    Middle Tested
  • Passed.
    Still correct at page sizes 1, 2 and 3
    Top Tested
  • Passed.
    New posts added at the top mid-scroll cause no repeats or gaps
    Top Tested
  • Took 4 min 37 s.
  • First words after 4 min 30 s.
  • Wrote 29,186 tokens, 27,978 of them thinking.
  • 59¢ at API prices.
All numbers.
Time to finish4 min 37 s
Time to first words4 min 30 s
Input tokens2
Cached tokens1,408
Output tokens29,186
Thinking tokens27,978
Words in the answer510
Price at API rates59¢
F02The search box shows results for an old queryLow 7/7Max 7/7Same

What it asks for. Pastes a TypeScript controller for a search-as-you-type box with two bugs: typing fast can leave results for an earlier query on screen, and the loading spinner disappears too early. The API client cannot cancel requests. It asks for the whole fixed file with the same exports, the spinner on exactly while the newest search is loading, and code that Node can run with its built-in TypeScript support.

Both settings passed 7 of 7 checks.

AI answer from Opus 5.5 · low effort · 2 Oct 2026

7 of 7

Here's the fixed file:

export interface SearchOptions {
  // The API client. It returns a Promise and can't be cancelled.
  fetchResults: (query: string) => Promise<string[]>;
  // Draws the list. Pass an error as the second argument to show the error state.
  render: (results: string[], error?: Error) => void;
  // Shows or hides the spinner.
  setLoading: (loading: boolean) => void;
}

export function createSearch({ fetchResults, render, setLoading }: SearchOptions) {
  // Increments on every input. Only the input holding the current
  // value may touch the list or the spinner.
  let latestRequestId = 0;

  return {
    // Called on every keystroke with the current text of the box.
    async onInput(raw: string): Promise<void> {
      const requestId = ++latestRequestId;
      const query = raw.trim();

      if (query === "") {
        // This input is now the newest, so any request still running is stale.
        setLoading(false);
        render([]);
        return;
      }

      setLoading(true);

      let results: string[];
      try {
        results = await fetchResults(query);
      } catch (err) {
        if (requestId !== latestRequestId) return; // a newer input took over
        setLoading(false);
        render([], err instanceof Error ? err : new Error(String(err)));
        return;
      }

      if (requestId !== latestRequestId) return; // a newer input took over
      setLoading(false);
      render(results);
    },
  };
}

What was wrong

  1. Stale results. Every call to onInput rendered whatever its own request returned, whenever that request finished. Requests can finish out of order. So if you typed ab and then abc, and the ab request came back last, its results overwrote the abc list. The empty-box case had the same problem: an older request finishing after the box was cleared would fill the list again.
  2. Spinner turning off too early. Every call ran setLoading(false) in its finally. When any older request finished, it hid the spinner, even though the newest request was still loading.

The fix

Each call takes a number from a counter that goes up on every input, including empty ones. After the await, the call checks whether it is still the newest. If it isn't, it does nothing: it doesn't render and it doesn't change the spinner. Older requests still run to the end, since we can't cancel them, but their results are ignored.

Only the newest call can turn the spinner off. It does that right as it shows its results or error. Because emptying the box also takes a new number, it turns the spinner off, clears the list, and makes every earlier request stale.

I also moved render(results) out of the try. Before, if render itself threw, that error was caught and shown as a search error.

The file uses only type annotations, so Node's type stripping can run it without a compile step.

  • Floor2/2, 100% passed
  • Middle2/2, 100% passed
  • Top3/3, 100% passed
7 of 7 checks passedfloor 2/2, middle 2/2, top 3/3
  • Passed.
    The file loads under type stripping and a plain search works
    Floor Tested
  • Passed.
    A failing search still shows the error and hides the spinner
    Floor Tested
  • Passed.
    Only the newest query's results are left on screen
    Middle Tested
  • Passed.
    The spinner stays on until the newest request finishes, then goes off
    Middle Tested
  • Passed.
    An old request that fails does not touch the newer results
    Top Tested
  • Passed.
    Emptying the box clears the list, and late answers don't bring it back
    Top Tested
  • Passed.
    If the newest search fails, the error is shown and stays
    Top Tested
  • Took 9.3 s.
  • First words after 1.5 s.
  • Wrote 944 tokens.
  • 2.6¢ at API prices.
All numbers.
Time to finish9.3 s
Time to first words1.5 s
Input tokens2
Cached tokens1,359
Output tokens944
Thinking tokens0
Words in the answer442
Price at API rates2.6¢

AI answer from Opus 5.5 · max effort · 2 Oct 2026

7 of 7

What was wrong

  • Stale results. Every call rendered its own response whenever it arrived. The requests can't be cancelled and can finish out of order. So if the request for ca came back after the one for cat, it overwrote the list. Clearing the box had the same problem: a request still in flight would put its results back into the empty list.
  • Spinner. Every call turned the spinner off in its finally. When several requests overlapped, the first one to finish hid the spinner while the newest one was still loading. Clearing the box never turned the spinner off at all, so it stayed on until the old request finished.

The fix

Every call to onInput takes a new id from a counter, including calls for an empty box. When its request finishes, a call renders and changes the spinner only if its id is still the latest. Otherwise a newer input has replaced it, so its results or error are ignored. The request still runs to the end, but nothing uses its outcome.

The empty-box path also takes an id, so a late response can't undo the clear. It now turns the spinner off as well.

The fix adds only plain JavaScript and type annotations, so it still runs with Node's built-in TypeScript support.

export interface SearchOptions {
  // The API client. It returns a Promise and can't be cancelled.
  fetchResults: (query: string) => Promise<string[]>;
  // Draws the list. Pass an error as the second argument to show the error state.
  render: (results: string[], error?: Error) => void;
  // Shows or hides the spinner.
  setLoading: (loading: boolean) => void;
}

export function createSearch({ fetchResults, render, setLoading }: SearchOptions) {
  // Every call to onInput takes a new id, including an empty input. Only the
  // call holding the latest id may update the list or the spinner. Requests
  // can't be cancelled, so older ones still run to the end, but whatever they
  // return or throw is ignored.
  let latestId = 0;

  return {
    // Called on every keystroke with the current text of the box.
    async onInput(raw: string): Promise<void> {
      const id = ++latestId;
      const query = raw.trim();

      if (query === "") {
        // Taking a new id above has already made any request in flight stale,
        // so it can't refill the list when it finishes.
        render([]);
        setLoading(false);
        return;
      }

      setLoading(true);
      try {
        const results = await fetchResults(query);
        if (id === latestId) render(results);
      } catch (err) {
        if (id === latestId) render([], err instanceof Error ? err : new Error(String(err)));
      } finally {
        // An older request finishing must not hide the spinner of a newer one.
        if (id === latestId) setLoading(false);
      }
    },
  };
}
  • Floor2/2, 100% passed
  • Middle2/2, 100% passed
  • Top3/3, 100% passed
7 of 7 checks passedfloor 2/2, middle 2/2, top 3/3
  • Passed.
    The file loads under type stripping and a plain search works
    Floor Tested
  • Passed.
    A failing search still shows the error and hides the spinner
    Floor Tested
  • Passed.
    Only the newest query's results are left on screen
    Middle Tested
  • Passed.
    The spinner stays on until the newest request finishes, then goes off
    Middle Tested
  • Passed.
    An old request that fails does not touch the newer results
    Top Tested
  • Passed.
    Emptying the box clears the list, and late answers don't bring it back
    Top Tested
  • Passed.
    If the newest search fails, the error is shown and stays
    Top Tested
  • Took 2 min 33 s.
  • First words after 2 min 28 s.
  • Wrote 16,694 tokens, 15,772 of them thinking.
  • 34¢ at API prices.
All numbers.
Time to finish2 min 33 s
Time to first words2 min 28 s
Input tokens2
Cached tokens1,359
Output tokens16,694
Thinking tokens15,772
Words in the answer455
Price at API rates34¢
F03Some subscriptions renew a day earlyLow 7/7Max 7/7Same

What it asks for. Pastes a JavaScript function that lists a subscription's monthly renewal dates, with its rules in a comment, and reports two symptoms: customers far from UTC renew a day early, and a customer who signed up on the last day of January has been charged on the 28th of every month since. It asks for the function fixed for Node 22 with no libraries, and a short note on what was wrong.

Both settings passed 7 of 7 checks.

AI answer from Opus 5.5 · low effort · 2 Oct 2026

7 of 7

Bugs: uses UTC date not local date; mutates d so day drifts (31→28 sticks). Fix: get local Y/M/D via Intl.DateTimeFormat, then compute pure calendar math.

// Monthly renewal dates for a subscription.
//
// signupIso  : the moment the customer signed up, as an ISO string in UTC,
//              e.g. "2025-01-31T23:30:00Z"
// timeZone   : the customer's IANA time zone, e.g. "Pacific/Auckland"
// count      : how many upcoming renewals to return
// returns    : an array of count local calendar dates, "YYYY-MM-DD", in the
//              customer's time zone, earliest first
//
// Rules:
//  - The first renewal is one month after the signup, not on the signup day.
//  - A renewal falls on the same day of the month as the signup's local date
//    in the customer's time zone.
//  - If a month is too short for that day, use its last day.
//  - Always count from the original signup day, never from the previous renewal.
export function renewalDates(signupIso, timeZone, count) {
  const signup = new Date(signupIso);
  if (Number.isNaN(signup.getTime())) {
    throw new RangeError(`Invalid signup date: ${signupIso}`);
  }

  // Signup's calendar date in the customer's time zone.
  const parts = new Intl.DateTimeFormat("en-US", {
    timeZone,
    year: "numeric",
    month: "numeric",
    day: "numeric",
  }).formatToParts(signup);
  const get = (type) => Number(parts.find((p) => p.type === type).value);
  const year = get("year");
  const month = get("month") - 1; // 0-based
  const day = get("day");

  const pad = (n, len = 2) => String(n).padStart(len, "0");
  const dates = [];

  for (let i = 1; i <= count; i++) {
    // Always offset from the original signup month, never the previous renewal.
    const total = month + i;
    const y = year + Math.floor(total / 12);
    const m = ((total % 12) + 12) % 12;
    // Day 0 of the next month = last day of month m.
    const lastDay = new Date(Date.UTC(y, m + 1, 0)).getUTCDate();
    const d = Math.min(day, lastDay);
    dates.push(`${pad(y, 4)}-${pad(m + 1)}-${pad(d)}`);
  }
  return dates;
}

What was wrong

  1. Wrong time zone. It read the day with getUTCDate(), so it used the signup's UTC date instead of the customer's local date. A morning signup on the 1st in Auckland is still the previous day in UTC (e.g. the 30th or 31st), so renewals landed a day early. The fix gets the local date with Intl.DateTimeFormat in the customer's timeZone.
  2. Counting from the previous renewal. It changed the same Date object on every pass and re-read the day from it. Once February clamped the 31st to the 28th, every later month started from 28. The fix keeps the original signup day and works out each renewal as signup month + i. It clamps only that month's result to the month's last day.

It also now does pure calendar arithmetic with no time-of-day involved, so DST changes can't move a date.

  • Floor2/2, 100% passed
  • Middle2/2, 100% passed
  • Top3/3, 100% passed
7 of 7 checks passedfloor 2/2, middle 2/2, top 3/3
  • Passed.
    Mid-month signups in UTC get the right dates
    Floor Tested
  • Passed.
    Returns exactly count dates as YYYY-MM-DD, in order
    Floor Tested
  • Passed.
    Uses the customer's local date (Auckland, Kolkata, LA, London)
    Middle Tested
  • Passed.
    Explains both causes: the UTC date and drifting month to month
    Middle Read
  • Passed.
    Signups on the 29th to 31st keep their day after short months
    Top Tested
  • Passed.
    Leap and non-leap Februaries are handled, including 2100
    Top Tested
  • Passed.
    Signups just before or after a clock change get the right local date
    Top Tested
  • Took 11 s.
  • First words after 2.1 s.
  • Wrote 1,127 tokens.
  • 2.9¢ at API prices.
All numbers.
Time to finish11 s
Time to first words2.1 s
Input tokens2
Cached tokens1,382
Output tokens1,127
Thinking tokens0
Words in the answer472
Price at API rates2.9¢

AI answer from Opus 5.5 · max effort · 2 Oct 2026

7 of 7

// Monthly renewal dates for a subscription.
//
// signupIso  : the moment the customer signed up, as an ISO string in UTC,
//              e.g. "2025-01-31T23:30:00Z"
// timeZone   : the customer's IANA time zone, e.g. "Pacific/Auckland"
// count      : how many upcoming renewals to return
// returns    : an array of count local calendar dates, "YYYY-MM-DD", in the
//              customer's time zone, earliest first
//
// Rules:
//  - The first renewal is one month after the signup, not on the signup day.
//  - A renewal falls on the same day of the month as the signup's local date
//    in the customer's time zone.
//  - If a month is too short for that day, use its last day.
//  - Always count from the original signup day, never from the previous renewal.
export function renewalDates(signupIso, timeZone, count) {
  // The signup's calendar date in the customer's own time zone. A fixed
  // locale keeps the parts as plain digits. Throws a RangeError for an
  // invalid timestamp or an unknown time zone.
  const parts = new Intl.DateTimeFormat('en-US', {
    timeZone,
    year: 'numeric',
    month: 'numeric',
    day: 'numeric',
  }).formatToParts(new Date(signupIso));
  const part = (type) => Number(parts.find((p) => p.type === type).value);
  const year = part('year');
  const monthIndex = part('month') - 1; // 0-11, as Date.UTC expects
  const day = part('day');

  const dates = [];
  for (let n = 1; n <= count; n++) {
    // Renewal n is n months after the signup month, on the signup's day.
    // It never uses the previous renewal. Date.UTC carries month indexes
    // past 11 over into the following years.
    const lastDay = new Date(Date.UTC(year, monthIndex + n + 1, 0)).getUTCDate();
    const renewal = new Date(Date.UTC(year, monthIndex + n, Math.min(day, lastDay)));
    dates.push(renewal.toISOString().slice(0, 10));
  }
  return dates;
}

What was wrong

  1. timeZone was never used. The signup day came from getUTCDate(), which gives the date in UTC, not the customer's local date. Auckland is 12–13 hours ahead of UTC. A signup there on the morning of the 1st is still the last day of the previous month in UTC, so the code used that day. The fix reads the signup's local date in the customer's zone with Intl.DateTimeFormat. The same bug works the other way west of UTC: an evening signup there is already the next day in UTC, so those customers have been renewed late. Customers rarely complain about that, so it won't show up in tickets.
  2. Each renewal was based on the previous one. day was read again from d on every pass of the loop. By then d held the previous renewal, already cut down to fit its month. Once February turned the 31st into the 28th, every later month got the 28th as well. Now every renewal is worked out from the original signup: renewal n is the signup month plus n, on the signup day, or on the last day of a shorter month.

One thing to check: the old code ignored timeZone, so a caller that never passed it would have gone unnoticed. If it arrives as undefined, Intl quietly uses the server's own time zone instead.

  • Floor2/2, 100% passed
  • Middle2/2, 100% passed
  • Top3/3, 100% passed
7 of 7 checks passedfloor 2/2, middle 2/2, top 3/3
  • Passed.
    Mid-month signups in UTC get the right dates
    Floor Tested
  • Passed.
    Returns exactly count dates as YYYY-MM-DD, in order
    Floor Tested
  • Passed.
    Uses the customer's local date (Auckland, Kolkata, LA, London)
    Middle Tested
  • Passed.
    Explains both causes: the UTC date and drifting month to month
    Middle Read
  • Passed.
    Signups on the 29th to 31st keep their day after short months
    Top Tested
  • Passed.
    Leap and non-leap Februaries are handled, including 2100
    Top Tested
  • Passed.
    Signups just before or after a clock change get the right local date
    Top Tested
  • Took 2 min 46 s.
  • First words after 2 min 39 s.
  • Wrote 18,545 tokens, 17,385 of them thinking.
  • 38¢ at API prices.
All numbers.
Time to finish2 min 46 s
Time to first words2 min 39 s
Input tokens2
Cached tokens1,382
Output tokens18,545
Thinking tokens17,385
Words in the answer522
Price at API rates38¢
F04Review a pull request before mergingLow 6/7Max 7/7Max +1

What it asks for. Pastes a pull request of about 170 lines that adds invoice endpoints to an Express and TypeScript API, and asks for a review before merging that lists what needs fixing in order of importance plus other points worth noting, in under 400 words.

Max effort passed 1 more check than low.

AI answer from Opus 5.5 · low effort · 2 Oct 2026

6 of 7

Must fix before merging

1. Any logged-in user can read any invoice (IDOR). GET /invoices/:id calls findInvoiceById(id) with no account check, so anyone with a valid session can fetch another account's invoice by guessing or enumerating IDs. Scope the query to the account:

export async function findInvoiceById(id: string, accountId: string) {
  const { rows } = await db.query(
    `SELECT ${COLUMNS} FROM invoices WHERE id = $1 AND account_id = $2`,
    [id, accountId],
  );
  return rows[0] ?? null;
}

Keep returning 404 rather than 403 when the account doesn't match, so the endpoint doesn't reveal which IDs exist.

2. to is not inclusive, despite what the code comment and docs say. created_at < $3 with to=2026-01-31 compares against midnight at the start of the 31st, so that whole day is excluded. Use created_at < ($3::date + 1) or add one day in the route.

3. Invalid dates cause 500s. The regex accepts values like 2026-02-30 and 2026-13-99. Postgres rejects them, and the handler surfaces that as an error instead of a 400. Check that the value is a real calendar date before querying.

4. Async errors may not be handled. If you're on Express 4, a rejected promise in these handlers (DB down, bad input) is never passed to your error middleware. The request hangs and you get an unhandled rejection. Wrap the handlers, or follow whatever pattern receipts.ts uses. Express 5 handles this for you.

Worth checking

  • The fmtMoney → formatMoney rename is unrelated to invoices and breaks any caller this diff doesn't touch. Grep for fmtMoney (for example in account.ts) and confirm tsc passes. Consider splitting the rename into its own PR.
  • Time zones: "day" boundaries depend on the column type (timestamp vs timestamptz) and the session time zone. Decide whose midnight counts (UTC or the account's) and document it.
  • No pagination or range limit: a very wide from/to returns everything in one response. Consider adding limit/cursor support or capping the range.
  • Unstable ordering: sorting by amount or created_at alone gives an arbitrary order for ties. Add , id as a tiebreaker, which you'll need anyway for pagination.
  • Silent fallbacks: an invalid sort/dir quietly uses the defaults, and from > to silently returns an empty list. Returning 400 would be friendlier, but this is your call.
  • formatMoney assumes two decimal places, which is wrong for currencies like JPY or KWD. This predates the PR but now affects more endpoints.

Looks good

  • The sort whitelist using Object.hasOwn and the hard-coded direction is safe from SQL injection.
  • The date and account filters are properly parameterized.
  • requireLogin covers the whole router.
  • The docs match the intended behaviour, once fix 2 makes the code match them.
  • Floor1/2, 50% passed
  • Middle2/2, 100% passed
  • Top3/3, 100% passed
6 of 7 checks passedfloor 1/2, middle 2/2, top 3/3
  • Passed.
    Gives a clear verdict: do not merge as it is
    Floor Counted
  • Missed.
    Stays under the 400-word limit429 words
    Floor Counted
  • Passed.
    Spots that any signed-in user can read another account's invoice
    Middle Counted
  • Passed.
    Gives a concrete fix: scope the lookup by account or return 404/403
    Middle Counted
  • Passed.
    Spots that the end date is left out although the docs say inclusive
    Top Counted
  • Passed.
    Puts the invoice access hole first, ahead of the other issues
    Top Read
  • Passed.
    Does not call the ORDER BY string building an injection risk
    Top Read
  • Took 16 s.
  • First words after 6.2 s.
  • Wrote 1,510 tokens, 428 of them thinking.
  • 4.6¢ at API prices.
All numbers.
Time to finish16 s
Time to first words6.2 s
Input tokens2
Cached tokens3,191
Output tokens1,510
Thinking tokens428
Words in the answer458
Price at API rates4.6¢

AI answer from Opus 5.5 · max effort · 2 Oct 2026

7 of 7

Must fix before merging

  1. GET /invoices/:id exposes other accounts' invoices. findInvoiceById looks up by id only, so any logged-in user can read any invoice whose id they know or guess. Scope the query with WHERE id = $1 AND account_id = $2, passing req.user.accountId. Another account's invoice then gets the same 404 as a missing one, so nothing leaks. Also make accountId a required parameter of the repo function so this can't recur.
  2. to isn't inclusive. The docs and route comment say it is, but created_at < $3 with a bare day excludes the whole last day. to=2026-01-31 drops Jan 31, and a single-day query (from = to) always returns nothing. Use created_at < $3::date + 1.
  3. Bad input causes uncaught database errors. The regex accepts impossible dates like 2026-02-30 or 2026-13-01, which Postgres rejects. A malformed :id fails the same way if id is a uuid column. Nothing catches these errors:
    • Express 4 without an async wrapper: the rejection is unhandled, so the request hangs and the process can crash on current Node. Any logged-in user can trigger this.
    • Express 5: the user gets a 500 instead of a 400.
    Fix it in three places: check that dates are real (e.g. the string round-trips through Date), check the id format, and make sure async errors reach your error handler.

Also worth knowing

  • The fmtMoney → formatMoney rename only updates receipts.ts. Search for other callers (e.g. the account routes) before merging. This unrelated refactor could also go in its own PR.
  • There's no limit or pagination, so a wide date range returns everything.
  • If created_at is timestamptz, day boundaries follow the database session's time zone. Decide what a "day" means (probably UTC) and document it.
  • sort=amount compares raw cents across currencies, which is meaningless if an account can have more than one. Also add id as a tiebreaker so the order is stable.
  • Unknown sort/dir values silently fall back to the defaults. Document that, plus that from/to are required, plus the 400/404 responses.
  • created_at is typed string, but node-postgres returns timestamps as Date unless you've changed its type parsers.
  • Add tests for the cross-account 404, the inclusive to boundary and invalid dates.

What's fine: the sort/direction whitelist is safe from SQL injection, and the list query is correctly scoped to the account.

  • Floor2/2, 100% passed
  • Middle2/2, 100% passed
  • Top3/3, 100% passed
7 of 7 checks passedfloor 2/2, middle 2/2, top 3/3
  • Passed.
    Gives a clear verdict: do not merge as it is
    Floor Counted
  • Passed.
    Stays under the 400-word limit
    Floor Counted
  • Passed.
    Spots that any signed-in user can read another account's invoice
    Middle Counted
  • Passed.
    Gives a concrete fix: scope the lookup by account or return 404/403
    Middle Counted
  • Passed.
    Spots that the end date is left out although the docs say inclusive
    Top Counted
  • Passed.
    Puts the invoice access hole first, ahead of the other issues
    Top Read
  • Passed.
    Does not call the ORDER BY string building an injection risk
    Top Read
  • Took 3 min 31 s.
  • First words after 3 min 26 s.
  • Wrote 25,038 tokens, 24,149 of them thinking.
  • 52¢ at API prices.
All numbers.
Time to finish3 min 31 s
Time to first words3 min 26 s
Input tokens2
Cached tokens3,191
Output tokens25,038
Thinking tokens24,149
Words in the answer394
Price at API rates52¢
F05SQL: which customers left and came backLow 7/7Max 7/7Same

What it asks for. Gives the schema of a subscriptions database and asks for one SQLite query that lists each gap in a customer's access that ended with them returning, with the last day of access, the day it came back and the number of days without it. Plan changes can make periods touch, overlap or sit inside one another, and none of that counts as a gap. A small example with its expected output is included.

Both settings passed 7 of 7 checks.

AI answer from Opus 5.5 · low effort · 2 Oct 2026

7 of 7

Query

WITH ordered AS (
  SELECT
    sp.customer_id,
    sp.start_date,
    -- Latest access day covered by any EARLIER period of this customer.
    -- NULL end_date (still active) is treated as "covers everything after".
    MAX(COALESCE(sp.end_date, '9999-12-31')) OVER (
      PARTITION BY sp.customer_id
      ORDER BY sp.start_date, sp.id
      ROWS BETWEEN UNBOUNDED PRECEDING AND 1 PRECEDING
    ) AS prev_covered_until
  FROM subscription_periods AS sp
)
SELECT
  c.name                                         AS name,
  o.prev_covered_until                           AS last_access,
  o.start_date                                   AS came_back,
  CAST(julianday(o.start_date)
       - julianday(o.prev_covered_until) - 1 AS INTEGER) AS days_without
FROM ordered AS o
JOIN customers AS c ON c.id = o.customer_id
WHERE o.start_date > date(o.prev_covered_until, '+1 day')
ORDER BY c.name, came_back, c.id;

How it works

  1. Sort each customer's periods by date. The window orders by start_date, so the order of the ids doesn't matter. id is only used to break ties between periods that start on the same day.
  2. Track how far access already reaches. For each period, prev_covered_until is the latest end_date among all periods that started before it. Using a running MAX instead of just the previous row's end date matters in two cases:
    • A short period that sits inside a longer one doesn't hide the longer one's end date.
    • Overlapping periods are handled correctly.
    An active period (NULL end date) counts as '9999-12-31'. Because of that, nothing that comes after an active period can ever look like a gap.
  3. Find the gaps. A period starts a gap only if it begins after the day following prev_covered_until. Periods that touch, start on the same day, or overlap all fail that test, so they aren't reported.
    • A customer's first period has no earlier periods, so prev_covered_until is NULL. The comparison is then NULL and the row is excluded.
    • A customer who left and never came back has no later period, so they produce no row.
  4. Count the days. julianday(came_back) - julianday(last_access) - 1 gives the number of whole days without access. For 2024-06-30 → 2024-07-10 that is 9.
  5. Sort. ORDER BY name uses SQLite's default BINARY collation, which is plain alphabetical order. After that, rows are sorted by came_back, then by customer id.

On your sample data, the query returns exactly the three expected rows. Bruno's overlapping periods (02-01 to 02-29, then 02-20 onward) don't count as a gap.

  • Floor2/2, 100% passed
  • Middle2/2, 100% passed
  • Top3/3, 100% passed
7 of 7 checks passedfloor 2/2, middle 2/2, top 3/3
  • Passed.
    The query runs and returns the four named columns
    Floor Tested
  • Passed.
    It gets the example in the question right
    Floor Tested
  • Passed.
    Periods that touch or overlap are not counted as gaps
    Middle Tested
  • Passed.
    Several gaps per customer, same-name customers, sort order
    Middle Tested
  • Passed.
    A period inside a longer one does not hide or invent a gap
    Top Tested
  • Passed.
    Open-ended (NULL end) periods never produce a gap after them
    Top Tested
  • Passed.
    Days without access is exact across month, leap and year ends
    Top Tested
  • Took 14 s.
  • First words after 6.2 s.
  • Wrote 1,451 tokens, 462 of them thinking.
  • 3.7¢ at API prices.
All numbers.
Time to finish14 s
Time to first words6.2 s
Input tokens2
Cached tokens1,682
Output tokens1,451
Thinking tokens462
Words in the answer380
Price at API rates3.7¢

AI answer from Opus 5.5 · max effort · 2 Oct 2026

7 of 7

WITH runs AS (
  SELECT
    customer_id,
    start_date,
    -- latest day with access among all periods that started before this one
    MAX(COALESCE(end_date, '9999-12-31')) OVER (
      PARTITION BY customer_id
      ORDER BY start_date, end_date, id
      ROWS BETWEEN UNBOUNDED PRECEDING AND 1 PRECEDING
    ) AS covered_until
  FROM subscription_periods
  WHERE end_date IS NULL OR end_date >= start_date
)
SELECT
  c.name,
  r.covered_until AS last_access,
  r.start_date    AS came_back,
  CAST(julianday(r.start_date) - julianday(r.covered_until) AS INTEGER) - 1 AS days_without
FROM runs AS r
JOIN customers AS c ON c.id = r.customer_id
WHERE julianday(r.start_date) - julianday(r.covered_until) > 1
ORDER BY c.name, r.start_date, c.id;

On your sample data this returns exactly the three rows you listed.

How it works

  • It tracks coverage so far, not the previous row. Each customer's periods are sorted by start_date, and the ids are only used to break ties. For each period, covered_until is the latest last-access day of any period that started earlier. It is a running MAX over the preceding rows. A plain LAG(end_date) would fail when a short period sits inside a longer one. The running max handles touching, same-day, overlapping and nested periods the same way.
  • Active periods count as never ending. A NULL end date is treated as '9999-12-31', so no gap can appear after an active period.
  • A gap is a period that starts more than one day after covered_until. last_access is covered_until and came_back is the new start date. days_without is the difference in days minus one: 06-30 to 07-10 is 10 days apart, so 9 days without access. Touching periods give 0 and overlapping ones give a negative number, so both are filtered out. The CAST makes the result an integer rather than 9.0.
  • No rows appear before the first period or after the last one. A customer's first period has nothing before it, so covered_until is NULL and the row is dropped. Gaps are only found when a later period starts, so a customer who left for good gets no row for that final departure.
  • Two details are defensive. Two periods starting on the same day are harmless, because the second one always sees the first one's coverage. The WHERE end_date >= start_date filter drops malformed periods that end before they start, since they never gave any access.

ORDER BY c.name uses SQLite's default byte-order comparison, which is case-sensitive. If some names start with a lowercase letter, use c.name COLLATE NOCASE instead.

  • Floor2/2, 100% passed
  • Middle2/2, 100% passed
  • Top3/3, 100% passed
7 of 7 checks passedfloor 2/2, middle 2/2, top 3/3
  • Passed.
    The query runs and returns the four named columns
    Floor Tested
  • Passed.
    It gets the example in the question right
    Floor Tested
  • Passed.
    Periods that touch or overlap are not counted as gaps
    Middle Tested
  • Passed.
    Several gaps per customer, same-name customers, sort order
    Middle Tested
  • Passed.
    A period inside a longer one does not hide or invent a gap
    Top Tested
  • Passed.
    Open-ended (NULL end) periods never produce a gap after them
    Top Tested
  • Passed.
    Days without access is exact across month, leap and year ends
    Top Tested
  • Took 2 min 20 s.
  • First words after 2 min 13 s.
  • Wrote 15,148 tokens, 14,188 of them thinking.
  • 31¢ at API prices.
All numbers.
Time to finish2 min 20 s
Time to first words2 min 13 s
Input tokens2
Cached tokens1,682
Output tokens15,148
Thinking tokens14,188
Words in the answer402
Price at API rates31¢
F06A contacts import that survives real CSV filesLow 9/9Max 0/9Max −9

What it asks for. Asks for a Python function, standard library only, that imports contacts from CSV files exported by many different tools and returns each contact's name, email and phone. The files vary in header names, comma or semicolon separators, a byte order mark, quoted fields with line breaks, Windows line endings and older files in Windows-1252. Rows without an email are skipped, and a repeated email keeps its first row.

Max effort passed 9 fewer checks than low.

AI answer from Opus 5.5 · low effort · 2 Oct 2026

9 of 9

import csv
import io

_NAME_HEADERS = {"name", "full name", "contact name"}
_EMAIL_HEADERS = {"email", "e-mail", "email address", "e-mail address"}
_PHONE_HEADERS = {"phone", "phone number", "mobile", "mobile phone", "telephone"}


def _decode(data):
    if data.startswith(b"\xef\xbb\xbf"):
        data = data[3:]
    try:
        return data.decode("utf-8")
    except UnicodeDecodeError:
        # Old Windows tool. A few bytes are undefined in cp1252, so replace them.
        return data.decode("cp1252", errors="replace")


def _detect_delimiter(text):
    # Count separators outside quotes on the first non-blank logical line (the header).
    commas = semicolons = 0
    in_quotes = False
    seen_content = False
    for ch in text:
        if ch == '"':
            in_quotes = not in_quotes
        elif not in_quotes:
            if ch == "\n":
                if seen_content:
                    break
                continue
            if ch == ",":
                commas += 1
            elif ch == ";":
                semicolons += 1
            if not ch.isspace():
                seen_content = True
    return ";" if semicolons > commas else ","


def _norm_header(h):
    return " ".join(h.strip().lower().split())


def import_contacts(path):
    with open(path, "rb") as f:
        text = _decode(f.read())

    # Normalise line endings so line breaks inside quoted fields become "\n".
    text = text.replace("\r\n", "\n").replace("\r", "\n")
    if not text.strip():
        return []

    reader = csv.reader(io.StringIO(text, newline=""),
                        delimiter=_detect_delimiter(text))

    # The header is the first row that is not blank.
    header = None
    for row in reader:
        if any(cell.strip() for cell in row):
            header = row
            break
    if header is None:
        return []

    name_idx = email_idx = phone_idx = None
    for i, h in enumerate(header):
        key = _norm_header(h)
        if key in _NAME_HEADERS and name_idx is None:
            name_idx = i
        elif key in _EMAIL_HEADERS and email_idx is None:
            email_idx = i
        elif key in _PHONE_HEADERS and phone_idx is None:
            phone_idx = i
    if email_idx is None:
        return []

    def cell(row, idx):
        if idx is None or idx >= len(row):
            return ""
        return row[idx].strip()

    contacts = []
    seen = set()
    for row in reader:
        email = cell(row, email_idx).lower()
        if not email or email in seen:
            continue
        seen.add(email)
        contacts.append({
            "name": cell(row, name_idx),
            "email": email,
            "phone": cell(row, phone_idx),
        })
    return contacts

How it works

  • Decoding: The function reads the file as bytes and removes a UTF-8 BOM if there is one. It then decodes the text as UTF-8 and falls back to Windows-1252 if that fails.
  • Line endings: All line endings are turned into \n, so line breaks inside quoted fields come out as a single \n.
  • Separator: It counts commas and semicolons in the header line, ignoring any inside quotes, and uses whichever appears more often.
  • Columns: csv.reader handles quoting, doubled quotes and multi-line fields. The header names are lower-cased and their spacing is normalised, then matched against the accepted names to find each column's position.
  • Rows: Each data row is read by position, so short rows give empty strings and extra cells are ignored. Rows with no email, including blank rows and rows of only separators, are skipped. A set of emails already seen drops duplicates and keeps the first one, in file order.
  • Missing email column: If the file has no email column, the function returns an empty list.
  • Floor2/2, 100% passed
  • Middle3/3, 100% passed
  • Top4/4, 100% passed
9 of 9 checks passedfloor 2/2, middle 3/3, top 4/4
  • Passed.
    A plain comma-separated file is imported
    Floor Tested
  • Passed.
    Fields are stripped, emails lower-cased, rows without an email skipped
    Floor Tested
  • Passed.
    Header names vary in case, spacing, wording and order
    Middle Tested
  • Passed.
    A UTF-8 BOM and Windows line endings are handled
    Middle Tested
  • Passed.
    Quoted commas and doubled quotes are read correctly
    Middle Tested
  • Passed.
    Line breaks inside quoted fields are kept as one \n
    Top Tested
  • Passed.
    Semicolon files are read, even with commas in fields
    Top Tested
  • Passed.
    Windows-1252 files decode correctly and UTF-8 stays intact
    Top Tested
  • Passed.
    Duplicates, blank and short rows, and a file with every quirk at once
    Top Tested
  • Took 18 s.
  • First words after 7.9 s.
  • Wrote 1,874 tokens, 521 of them thinking.
  • 4.4¢ at API prices.
All numbers.
Time to finish18 s
Time to first words7.9 s
Input tokens2
Cached tokens1,394
Output tokens1,874
Thinking tokens521
Words in the answer493
Price at API rates4.4¢

AI answer from Opus 5.5 · max effort · 2 Oct 2026

0 of 9

The answer was cut off at the token limit.

  • Floor0/2, 0% passed
  • Middle0/3, 0% passed
  • Top0/4, 0% passed
0 of 9 checks passedfloor 0/2, middle 0/3, top 0/4
  • Missed.
    A plain comma-separated file is importedNo answer
    Floor Tested
  • Missed.
    Fields are stripped, emails lower-cased, rows without an email skippedNo answer
    Floor Tested
  • Missed.
    Header names vary in case, spacing, wording and orderNo answer
    Middle Tested
  • Missed.
    A UTF-8 BOM and Windows line endings are handledNo answer
    Middle Tested
  • Missed.
    Quoted commas and doubled quotes are read correctlyNo answer
    Middle Tested
  • Missed.
    Line breaks inside quoted fields are kept as one \nNo answer
    Top Tested
  • Missed.
    Semicolon files are read, even with commas in fieldsNo answer
    Top Tested
  • Missed.
    Windows-1252 files decode correctly and UTF-8 stays intactNo answer
    Top Tested
  • Missed.
    Duplicates, blank and short rows, and a file with every quirk at onceNo answer
    Top Tested
  • Took 4 min 55 s.
  • Wrote 32,000 tokens, 32,000 of them thinking.
  • 65¢ at API prices.
All numbers.
Time to finish4 min 55 s
Time to first wordsnot recorded
Input tokens2
Cached tokens1,394
Output tokens32,000
Thinking tokens32,000
Words in the answer0
Price at API rates65¢
F07Make a slow script fastLow 7/7Max 0/7Max −7

What it asks for. Pastes a Python function that matches bank payments to open invoices by customer, amount and a due date within 7 days, and says it is far too slow on about 200,000 of each. It asks for a replacement that runs in a couple of seconds, returns exactly the same results for every input and leaves the lists it is given unchanged.

Max effort passed 7 fewer checks than low.

AI answer from Opus 5.5 · low effort · 2 Oct 2026

7 of 7

The old function was slow because of a nested loop. For each of the 200,000 payments it scanned all 200,000 invoices, which is about 40 billion checks. The new version builds an index once and answers each payment with a short lookup.

from bisect import bisect_left
from collections import defaultdict


def match_payments(payments, invoices):
    """Match bank payments to open invoices.

    payments: list of dicts with id, customer, amount (integer cents) and date (datetime.date).
    invoices: list of dicts with id, customer, amount (integer cents) and due (datetime.date).

    Payments are handled one at a time, in the order they are given. A payment is
    matched to the unpaid invoice that has the same customer and the same amount,
    whose due date is within 7 days either side of the payment date (7 days
    apart still counts). If several invoices qualify, the one with the earliest
    due date wins; if they are due on the same day, the lowest invoice id wins.
    An invoice can be matched to one payment only.

    Returns (matches, unmatched):
      matches   - list of (payment_id, invoice_id), in payment order
      unmatched - list of ids of the payments that found no invoice, in payment order
    """
    # Group invoices by (customer, amount). The input lists are never modified.
    grouped = defaultdict(list)
    for inv in invoices:
        grouped[(inv["customer"], inv["amount"])].append((inv["due"], inv["id"]))

    # Within each group, sort by (due, id): the first eligible entry is the winner.
    # Each group holds: due ordinals (for bisect), ids, and a "next available"
    # pointer array used to skip over invoices that are already taken.
    index = {}
    for key, items in grouped.items():
        items.sort(key=lambda t: (t[0], t[1]))
        dues = [d.toordinal() for d, _ in items]
        ids = [i for _, i in items]
        nxt = list(range(len(items) + 1))  # nxt[len] is a sentinel
        index[key] = (dues, ids, nxt)

    def find(nxt, i):
        # Next index >= i that has not been removed (with path compression).
        root = i
        while nxt[root] != root:
            root = nxt[root]
        while nxt[i] != root:
            nxt[i], i = root, nxt[i]
        return root

    used = set()
    matches = []
    unmatched = []

    for pay in payments:
        group = index.get((pay["customer"], pay["amount"]))
        best_id = None
        found = False
        if group is not None:
            dues, ids, nxt = group
            n = len(dues)
            day = pay["date"].toordinal()
            hi = day + 7
            i = find(nxt, bisect_left(dues, day - 7))
            while i < n and dues[i] <= hi:
                inv_id = ids[i]
                if inv_id in used:
                    # Already used (possibly via a duplicate id elsewhere):
                    # remove it permanently and move on.
                    nxt[i] = i + 1
                    i = find(nxt, i + 1)
                    continue
                best_id = inv_id
                found = True
                nxt[i] = i + 1  # remove this invoice from the group
                break
        if not found:
            unmatched.append(pay["id"])
        else:
            used.add(best_id)
            matches.append((pay["id"], best_id))

    return matches, unmatched

What changed

  1. Grouping. Invoices are grouped once by (customer, amount). A payment only ever looks at invoices that could match it, instead of all 200,000.
  2. Sorting plus binary search. Each group is sorted by (due, id), which is exactly the old tie-break order. A binary search jumps straight to the first invoice due on or after payment date − 7. Scanning forward from there, the first available invoice whose due date is at most payment date + 7 is the one the old code would have picked.
  3. Skipping used invoices. Your large customer with over a thousand same-amount invoices would otherwise make each payment walk past every invoice already used. To avoid that, a used invoice is removed with a "next available" pointer array, compressed as it is followed. Each invoice is skipped at most a handful of times in total, not once per payment.

Matching the old behaviour exactly

  • Payment order. Payments are still processed in the order given, so unsorted bank order gives the same results.
  • Used invoices. "Used" is still tracked by invoice id in one global set, as before. If the same id appears more than once, even under a different customer or amount, it becomes unavailable everywhere once matched.
  • The 7-day limit. It is inclusive on both sides, as before. Day ordinals are used, which gives the same result as (due - date).days for datetime.date values.
  • Your two lists. payments and invoices are only read, never modified.

Speed

Building the index takes O(N log N) time, and each payment costs about O(log N) on top of that. 200,000 × 200,000 should run in roughly a second.

  • Floor2/2, 100% passed
  • Middle3/3, 100% passed
  • Top2/2, 100% passed
7 of 7 checks passedfloor 2/2, middle 3/3, top 2/2
  • Passed.
    Same results as the original on simple cases
    Floor Tested
  • Passed.
    Same results as the original on many random small inputs
    Floor Tested
  • Passed.
    The 7-day limit is identical at both edges
    Middle Tested
  • Passed.
    The two input lists are left unchanged
    Middle Tested
  • Passed.
    Fast on 20,000 payments and invoices, same results
    Middle Tested
  • Passed.
    Earliest due date, then lowest id, in payment order, each invoice once
    Top Tested
  • Passed.
    Under 6 seconds on 200,000 payments and invoices, same results
    Top Tested
  • Took 23 s.
  • First words after 10 s.
  • Wrote 2,438 tokens, 745 of them thinking.
  • 5.7¢ at API prices.
All numbers.
Time to finish23 s
Time to first words10 s
Input tokens2
Cached tokens1,561
Output tokens2,438
Thinking tokens745
Words in the answer741
Price at API rates5.7¢

AI answer from Opus 5.5 · max effort · 2 Oct 2026

0 of 7

The answer was cut off at the token limit.

  • Floor0/2, 0% passed
  • Middle0/3, 0% passed
  • Top0/2, 0% passed
0 of 7 checks passedfloor 0/2, middle 0/3, top 0/2
  • Missed.
    Same results as the original on simple casesNo answer
    Floor Tested
  • Missed.
    Same results as the original on many random small inputsNo answer
    Floor Tested
  • Missed.
    The 7-day limit is identical at both edgesNo answer
    Middle Tested
  • Missed.
    The two input lists are left unchangedNo answer
    Middle Tested
  • Missed.
    Fast on 20,000 payments and invoices, same resultsNo answer
    Middle Tested
  • Missed.
    Earliest due date, then lowest id, in payment order, each invoice onceNo answer
    Top Tested
  • Missed.
    Under 6 seconds on 200,000 payments and invoices, same resultsNo answer
    Top Tested
  • Took 5 min 38 s.
  • Wrote 32,000 tokens, 32,000 of them thinking.
  • 65¢ at API prices.
All numbers.
Time to finish5 min 38 s
Time to first wordsnot recorded
Input tokens2
Cached tokens1,561
Output tokens32,000
Thinking tokens32,000
Words in the answer0
Price at API rates65¢

Planning and architecture

A01Plan a booking app for a dog groomerLow 7/9Max 0/9Max −7

What it asks for. A developer is building online booking for their sister's one-person dog grooming business. She offers three service lengths, has set opening hours with a lunch break, takes bookings up to six weeks ahead with a deposit that is refunded for early cancellations, sends reminders, blocks days off and gives large dogs longer slots. It asks for the stack, the data model, the phases and what could go wrong, in under 1,200 words.

Max effort passed 7 fewer checks than low.

AI answer from Opus 5.5 · low effort · 2 Oct 2026

7 of 9

Online Booking for a One-Person Dog Groomer: The Plan

First, consider buying instead of building

Square Appointments, Acuity, Fresha and similar tools already handle deposits, reminders and blocked days, often for little or nothing at one-staff scale. Spend an hour checking whether one covers the large-dog rule. Usually it can, as a separate "large dog" version of each service.

If you build anyway (control, learning, no monthly fee), here's the plan.

Stack

Choose boring, well-documented tools you already know.

  • Backend: Django with PostgreSQL. You get auth, an ORM, migrations and an admin for free. Postgres matters here because it can prevent double-bookings at the database level.
  • Frontend: Server-rendered Django templates with a little JavaScript (or htmx). There are two views:
    • a public booking flow,
    • a mobile-first dashboard for your sister.
    A SPA isn't worth it at this size.
  • Payments: Stripe Checkout for the deposit, Stripe webhooks to confirm payment, and the Stripe Refunds API for cancellations.
  • Email: Postmark or Resend. SMS through Twilio is optional, but reminders by text get read far more often.
  • Scheduled jobs: A cron job (the host's scheduler, or a management command every 15 minutes) for reminders and for expiring unpaid holds. You don't need Celery.
  • Hosting: Render, Railway or Fly.io with managed Postgres. Turn on automated backups.
  • Time handling: Python's zoneinfo, with her IANA zone (for example Europe/London) stored in settings.

Data model

Service

  • name, base_minutes (30/60/90), price, active

Booking

  • service (FK)
  • dog_size (standard/large)
  • duration_minutes: snapshot at booking time, base or base × 2
  • price and deposit_amount: snapshots, so later price changes don't rewrite history
  • start, end: timezone-aware, stored in UTC
  • client_name, email, phone, dog_name, breed, notes
  • status: pending_payment → confirmed → completed / cancelled / no_show, plus expired
  • stripe_checkout_session_id, stripe_payment_intent_id, refund_id
  • cancel_token: a random secret for the "manage my booking" link, so clients don't need accounts
  • hold_expires_at, reminder_sent_at, cancelled_at, created_at

WorkingHours

  • weekday, open_time, close_time
  • Store breaks as separate rows (Tue 9:00–12:30, Tue 13:30–17:00). Lunch then needs no special case, and she can change hours later without a code change.

BlockedPeriod

  • start, end, reason. A full day is just a range covering it.

Settings (a single row or config)

  • timezone, booking horizon (42 days), deposit percent (20), refund cutoff (24h), slot granularity (for example 30 minutes), optional buffer between appointments

Preventing double-bookings

Add a Postgres exclusion constraint on tstzrange(start, end) for bookings whose status is pending_payment or confirmed. Two people clicking the same slot at the same moment then can't both win, whatever your Python code does.

Core logic: availability

Write this as a pure function, available_starts(date, duration), and test it thoroughly before building any UI.

  1. If the date is in the past, beyond 6 weeks, or not Tuesday to Saturday, return nothing.
  2. Build the working intervals for that date in local time, then convert them to aware datetimes.
  3. Subtract blocked periods and active bookings.
  4. Return each start time, on the slot grid, where start + duration fits entirely inside one free interval. An appointment must never span lunch.

A large-dog 90-minute groom takes 180 minutes. It fits in the morning (3.5h) or the afternoon (3.5h), but only at a few start times. That's correct, but she should know.

Phases

Phase 0: Decisions with your sister (one evening). Get answers to:

  • Does a large dog also cost double? This affects the deposit.
  • What counts as "large"? Pick a weight threshold the client declares, and decide what she does if someone under-declares.
  • Does she want buffer time between dogs?
  • What slot grid should start times use?
  • What time should reminders go out?
  • Can clients reschedule, or only cancel and rebook?
  • What's the no-show policy?
  • What does she want written in the policy text clients agree to?

Phase 1: Availability engine plus tests. Cover these cases:

  • lunch boundaries,
  • 180-minute bookings,
  • blocked days,
  • the 6-week edge,
  • both DST transition dates.

Phase 2: Booking without payment, plus her dashboard. The dashboard needs:

  • a today/week list,
  • booking details with a tap-to-call phone link,
  • a "block time" form,
  • a new-booking notification to her (email or push).

Run this privately with fake bookings so she can try it on her phone.

Phase 3: Deposits. Use this flow:

  1. Create the booking as pending_payment with a 15-minute hold.
  2. Redirect to Stripe Checkout.
  3. Confirm only when the checkout.session.completed webhook arrives. Never rely on the success-page redirect.
  4. A cron job expires stale holds.
  5. Make the webhook handler idempotent, because Stripe retries.

Phase 4: Cancellation and refunds. Use the token link and compare now against start - 24h in absolute time.

  • More than 24h ahead: refund through the API and record the refund ID.
  • Within 24h: cancel and keep the deposit.

She also needs a "cancel and refund anyway" button for goodwill cases.

Phase 5: Reminders. The cron job finds confirmed bookings for tomorrow (local date) with no reminder_sent_at, sends the reminder, then sets the field. Setting the field makes the job safe to re-run.

Phase 6: Polish and soft launch. Add:

  • confirmation emails with an .ics attachment,
  • a privacy notice,
  • error monitoring (Sentry),
  • a backup restore test.

Launch to a handful of regular clients first.

What could go wrong

Time zones and DST

  • Store UTC and compute availability in local time. Never add timedelta(days=1) to an aware datetime to get "tomorrow 9:00". Build it from the local date and time instead.
  • The DST changes fall on Sundays and Mondays, when she's closed, which helps. But "24 hours before" a Tuesday booking after a change is not the same as "same time Monday". Decide which one the policy means, and write the policy to match.
  • Server cron runs in UTC. A fixed UTC reminder time drifts by an hour twice a year unless you compute it from local time.

Payments

  • Race between payment and expiry. A client pays at minute 16 after their hold has expired and someone else took the slot. Detect this in the webhook and refund automatically.
  • Stripe fees aren't returned on refunds. She loses a little on every refunded cancellation. She should know this up front.
  • Chargebacks on kept deposits. A clear policy checkbox at booking time is her evidence.
  • Test mode vs live keys. Check webhook secrets separately in each environment.

Business rules

  • Blocking a day that already has bookings. The block form must show the conflicts and offer "cancel and refund all", not hide them silently.
  • Bookings made less than 24h ahead are non-refundable from the moment they're made. Show that clearly before payment.
  • Bookings made after the reminder run for tomorrow either get no reminder or need one sent immediately. Pick one.
  • Large-dog honesty. A self-declared size can be gamed, and a mistake wrecks her afternoon. Let her change the size and duration from the dashboard.
  • Changing hours or prices later. The snapshots protect existing bookings, but new working hours could leave old bookings "outside hours". Flag them; don't delete them.

Operations

  • She's running it from her phone, so every admin action must work one-handed on a small screen. Test on her actual phone.
  • Single point of failure: you. Write a one-page "what to do if it breaks" note. Keep a manual fallback: she can still take phone bookings and enter them in the dashboard.
  • Spam and bot bookings. Pending holds expire, but add rate limiting or a CAPTCHA on the booking form.
  • Personal data (names, phones, emails). Use HTTPS, a short privacy notice, and minimal retention. Check local data-protection rules.
  • Email deliverability. Set SPF/DKIM on her domain, or reminders land in spam.

Rough timeline

At weekend pace, this is about 6–8 weekends to a soft launch:

  • Phases 1–2: about three weekends
  • Payments and refunds: about two weekends
  • Reminders, polish and testing: one to two weekends
  • Floor2/2, 100% passed
  • Middle2/4, 50% passed
  • Top3/3, 100% passed
7 of 9 checks passedfloor 2/2, middle 2/4, top 3/3
  • Passed.
    Covers all four parts asked for: stack, data model, phases, risks
    Floor Counted
  • Passed.
    The data model has bookings, services and clients or dogs
    Floor Counted
  • Missed.
    Stays under the 1,200-word limit1268 words
    Middle Counted
  • Passed.
    Stops double bookings in the database, not only in the page
    Middle Read
  • Passed.
    Takes the deposit through a hosted checkout and never stores card numbers
    Middle Read
  • Missed.
    The first phase is small and usable before payments existFirst build phase is "Availability engine plus tests" only; a working booking flow arrives in Phase 2, so Phase 1 produ…
    Middle Read
  • Passed.
    Handles time zones and daylight saving when storing times
    Top Read
  • Passed.
    Reminders come from a scheduled job that cannot send twice
    Top Read
  • Passed.
    Applies the 24-hour refund rule from the appointment start time
    Top Read
  • Took 44 s.
  • First words after 12 s.
  • Wrote 4,001 tokens, 855 of them thinking.
  • 8.5¢ at API prices.
All numbers.
Time to finish44 s
Time to first words12 s
Input tokens2
Cached tokens1,045
Output tokens4,001
Thinking tokens855
Words in the answer1,350
Price at API rates8.5¢

AI answer from Opus 5.5 · max effort · 2 Oct 2026

0 of 9

The answer was cut off at the token limit.

  • Floor0/2, 0% passed
  • Middle0/4, 0% passed
  • Top0/3, 0% passed
0 of 9 checks passedfloor 0/2, middle 0/4, top 0/3
  • Missed.
    Covers all four parts asked for: stack, data model, phases, risksNo answer
    Floor Counted
  • Missed.
    The data model has bookings, services and clients or dogsNo answer
    Floor Counted
  • Missed.
    Stays under the 1,200-word limitNo answer
    Middle Counted
  • Missed.
    Stops double bookings in the database, not only in the pageNo answer
    Middle Read
  • Missed.
    Takes the deposit through a hosted checkout and never stores card numbersNo answer
    Middle Read
  • Missed.
    The first phase is small and usable before payments existNo answer
    Middle Read
  • Missed.
    Handles time zones and daylight saving when storing timesNo answer
    Top Read
  • Missed.
    Reminders come from a scheduled job that cannot send twiceNo answer
    Top Read
  • Missed.
    Applies the 24-hour refund rule from the appointment start timeNo answer
    Top Read
  • Took 5 min 14 s.
  • Wrote 32,000 tokens, 32,000 of them thinking.
  • 65¢ at API prices.
All numbers.
Time to finish5 min 14 s
Time to first wordsnot recorded
Input tokens2
Cached tokens1,045
Output tokens32,000
Thinking tokens32,000
Words in the answer0
Price at API rates65¢
A02Should we switch to microservices?Low 7/8Max 8/8Max +1

What it asks for. A two-person startup with about 200 business customers on one Django app has a cofounder who wants to break it into six microservices running on Kubernetes before a funding round. Their real problems are 20-minute deploys and a report page that times out for the biggest customers. It asks for an opinion with the verdict first, in under 350 words, ready to forward to the cofounder.

Max effort passed 1 more check than low.

AI answer from Opus 5.5 · low effort · 2 Oct 2026

7 of 8

Don't split into microservices now. Neither of your real problems comes from the monolith, both can be fixed in days or weeks, and a six-service Kubernetes migration would take your whole two-person team for months while making deploys and debugging harder.

Here's why:

The report timeout is a query and data problem, not an architecture problem. Moving the report into its own service would give it the same slow queries plus network hops. The likely fixes:

  • Profile the report for the three big customers. Look for N+1 queries, missing indexes and unbounded loops in Python.
  • Add select_related/prefetch_related, push aggregation into SQL, and add indexes where EXPLAIN ANALYZE points.
  • Precompute monthly reports with a scheduled background job (Celery, or a cron plus a management command) or a materialized view. Then serve the stored result or email a link.

The 20-minute deploys are a pipeline problem. Find out where the time goes. It's usually some mix of:

  • Uncached dependency installs or Docker layers
  • A test suite running serially, which you can parallelize with pytest -n
  • Slow migrations or static asset builds

Six services means six pipelines, versioned APIs between them, distributed tracing, and a Kubernetes cluster for someone to run. With two engineers, that someone is you two, instead of building product.

On "that's how you scale": 200 business customers is a small load for one well-tuned Django and Postgres setup. Shopify, GitHub and Basecamp ran large monoliths for years. Microservices mainly solve an organizational problem, letting many teams deploy independently. You don't have that problem.

On the funding round: investors fund traction and speed. "We fixed our biggest customers' pain in two weeks and ship daily" is a better story than "we spent the quarter on infrastructure." Good technical diligence will see six services run by two people as a risk, not a strength.

What I'd do instead:

  1. Fix the report this week, starting with profiling and then precomputing.
  2. Get deploys under five minutes.
  3. Keep the monolith modular, with clean app boundaries and no cross-app reaching into models. That way, if you ever do need to extract a service, the seams are already there.

Revisit the question when you have several teams stepping on each other, or one component with clearly different scaling needs.

  • Floor1/2, 50% passed
  • Middle2/2, 100% passed
  • Top4/4, 100% passed
7 of 8 checks passedfloor 1/2, middle 2/2, top 4/4
  • Passed.
    Says up front: not now, keep the monolith
    Floor Counted
  • Missed.
    Stays under the 350-word limit376 words
    Floor Counted
  • Passed.
    Gives a concrete fix for the report timeout
    Middle Counted
  • Passed.
    Gives a concrete fix for the 20-minute deploys
    Middle Counted
  • Passed.
    Says when splitting would make sense
    Top Read
  • Passed.
    Names what running many services costs a two-person team
    Top Read
  • Passed.
    Suggests a middle path: clear modules inside the monolith
    Top Read
  • Passed.
    Tone is fair to the cofounder and safe to forward
    Top Read
  • Took 10 s.
  • First words after 1.7 s.
  • Wrote 772 tokens.
  • 2.0¢ at API prices.
All numbers.
Time to finish10 s
Time to first words1.7 s
Input tokens2
Cached tokens875
Output tokens772
Thinking tokens0
Words in the answer382
Price at API rates2.0¢

AI answer from Opus 5.5 · max effort · 2 Oct 2026

8 of 8

Don't split into microservices now. It would tie up both of you for months right before the funding round and fix neither of your real problems. Both the slow deploys and the report timeouts should be fixable inside the monolith in a few weeks.

Why microservices won't help

  • The timeout hits only your three biggest customers. That points to queries that slow down as data grows. A separate service would run the same queries on the same data, plus network hops.
  • Smaller services might deploy faster. Fixing the current pipeline gets you there without six pipelines, versioned APIs between services, and data split across databases.
  • Microservices mainly help teams scale, not traffic, by letting many teams ship independently. With two engineers you'd pay the costs (Kubernetes alone is a part-time job) without the benefit.
  • 200 business customers is a modest load for Django and Postgres. Instagram runs a Django monolith at vastly larger scale.

Fix the reports (1–2 weeks)

  1. Profile one big customer's report with EXPLAIN ANALYZE and Django Debug Toolbar. The usual causes are missing indexes, N+1 queries (fix with select_related/prefetch_related), or aggregating in Python instead of SQL.
  2. Generate reports in a background job (e.g. Celery) and notify the customer when the report is ready, so no web request times out.
  3. Completed months don't change, so build each report once and store it.

Fix deploys (about a week)

  1. Time each stage to see where the 20 minutes goes. It's usually tests, image builds or static assets.
  2. Run tests in parallel (pytest-xdist or manage.py test --parallel), and cache dependencies and build layers. Getting under five minutes is often achievable.

On the funding round Investors back revenue, growth and retention, not architecture, and a Django monolith is normal at your stage. "We fixed our biggest customers' problems in weeks" is a better story than "we spent months rewriting."

When to revisit: keep clean boundaries between Django apps now. Pull a piece out into its own service only when a real need appears, such as a separate team or a component that needs to scale very differently.

  • Floor2/2, 100% passed
  • Middle2/2, 100% passed
  • Top4/4, 100% passed
8 of 8 checks passedfloor 2/2, middle 2/2, top 4/4
  • Passed.
    Says up front: not now, keep the monolith
    Floor Counted
  • Passed.
    Stays under the 350-word limit
    Floor Counted
  • Passed.
    Gives a concrete fix for the report timeout
    Middle Counted
  • Passed.
    Gives a concrete fix for the 20-minute deploys
    Middle Counted
  • Passed.
    Says when splitting would make sense
    Top Read
  • Passed.
    Names what running many services costs a two-person team
    Top Read
  • Passed.
    Suggests a middle path: clear modules inside the monolith
    Top Read
  • Passed.
    Tone is fair to the cofounder and safe to forward
    Top Read
  • Took 2 min 57 s.
  • First words after 2 min 53 s.
  • Wrote 20,989 tokens, 20,272 of them thinking.
  • 42¢ at API prices.
All numbers.
Time to finish2 min 57 s
Time to first words2 min 53 s
Input tokens2
Cached tokens875
Output tokens20,989
Thinking tokens20,272
Words in the answer354
Price at API rates42¢
A03Rename a column with no downtimeLow 7/9Max 0/9Max −7

What it asks for. Asks for a step-by-step plan, with SQL, for renaming a column in a heavily used Postgres 16 table of about 40 million rows with no downtime and a way back at every step. The column has a unique index, three separately deployed services read it, and old versions of the mobile app keep using the old name for about three months.

Max effort passed 7 fewer checks than low.

AI answer from Opus 5.5 · low effort · 2 Oct 2026

7 of 9

Zero-downtime rename: users.username → users.handle

You can't use a plain ALTER TABLE ... RENAME COLUMN. The rename itself is instant, but every deployed reader and writer that uses the old name breaks at that moment. Because three services deploy separately, use expand → migrate → contract: run both columns side by side, keep them in sync with a trigger, move each service across, then remove the old column.

Treat the database column and the JSON field as two separate migrations. The JSON username field must stay for at least three months. The column can be gone long before that.

General rule for every DDL step: run it with a short lock timeout and retry on failure. A DDL statement waiting behind a long transaction will block all traffic queued behind it.

SET lock_timeout = '2s';
SET statement_timeout = 0;  -- per session, for the long-running steps

Phase 1: Expand (database only, no app changes)

1. Add the new column (nullable, no default)

ALTER TABLE users ADD COLUMN handle text;

This is a metadata-only change and takes milliseconds.

  • Rollback: ALTER TABLE users DROP COLUMN handle;

2. Add a sync trigger that works in both directions

Any write to either column is mirrored to the other. This lets old and new code run at the same time in any order.

CREATE OR REPLACE FUNCTION users_sync_handle() RETURNS trigger
LANGUAGE plpgsql AS $$
BEGIN
  IF TG_OP = 'INSERT' THEN
    NEW.handle   := COALESCE(NEW.handle, NEW.username);
    NEW.username := COALESCE(NEW.username, NEW.handle);
  ELSIF NEW.username IS DISTINCT FROM OLD.username
        AND NEW.handle IS NOT DISTINCT FROM OLD.handle THEN
    NEW.handle := NEW.username;          -- old code wrote username
  ELSIF NEW.handle IS DISTINCT FROM OLD.handle
        AND NEW.username IS NOT DISTINCT FROM OLD.username THEN
    NEW.username := NEW.handle;          -- new code wrote handle
  END IF;
  RETURN NEW;
END $$;

CREATE TRIGGER users_sync_handle
BEFORE INSERT OR UPDATE ON users
FOR EACH ROW EXECUTE FUNCTION users_sync_handle();

NOT NULL is checked after BEFORE triggers run. So once new code inserts only handle, the username NOT NULL constraint is still satisfied.

  • Rollback:
    DROP TRIGGER users_sync_handle ON users;
    DROP FUNCTION users_sync_handle();

3. Backfill in batches

Run this from a script. Each batch is its own transaction, with a short sleep between batches. Watch replication lag and autovacuum while it runs.

UPDATE users
SET    handle = username
WHERE  id >= $1 AND id < $2      -- e.g. 10k–50k ids per batch
AND    handle IS NULL;

Repeat until this returns 0:

SELECT count(*) FROM users WHERE handle IS NULL;
  • Rollback: nothing to undo. The data is only a copy. Stop the script, or drop the column as in step 1.

4. Add the unique index without blocking writes

CREATE UNIQUE INDEX CONCURRENTLY users_handle_key ON users (handle);

If it fails, it leaves an INVALID index behind. Drop that index and retry.

If your existing uniqueness is a constraint rather than just an index (check with \d users), recreate it the same way:

ALTER TABLE users ADD CONSTRAINT users_handle_key UNIQUE USING INDEX users_handle_key;
  • Rollback: DROP INDEX CONCURRENTLY users_handle_key;

5. Make handle NOT NULL without a long lock

Validating the check scans the table while holding only a light lock. On Postgres 12 and later, SET NOT NULL then uses the validated check and skips its own full scan.

ALTER TABLE users ADD CONSTRAINT users_handle_not_null
  CHECK (handle IS NOT NULL) NOT VALID;
ALTER TABLE users VALIDATE CONSTRAINT users_handle_not_null;
ALTER TABLE users ALTER COLUMN handle SET NOT NULL;
ALTER TABLE users DROP CONSTRAINT users_handle_not_null;
  • Rollback: ALTER TABLE users ALTER COLUMN handle DROP NOT NULL;

At this point both columns are always identical, unique and NOT NULL.


Phase 2: Migrate the applications

6. Move the two read-only services to handle

Deploy them independently, in any order. They are safe because the trigger keeps both columns identical.

  • Rollback: redeploy the previous version, which reads username. It still works.

7. Move the main API's reads to handle, and change its JSON output

  • Read from handle.
  • Return both fields in the JSON, with the same value:
    { "username": "alice", "handle": "alice" }
  • If clients send username in request bodies, accept both, and prefer handle when both are present.

Release the mobile app version that uses handle. Its three-month clock starts now.

  • Rollback: redeploy the previous API version.

8. Move the main API's writes to handle only

Stop writing username. The trigger fills it in, so any old instance or reader still using username keeps seeing correct data.

  • Rollback: redeploy the previous version. The trigger handles both write paths, so either version is correct.

9. Verify that nothing touches username any more

  • Search all three codebases, including raw SQL, ORM models, reports, ETL jobs, BI tools and ad-hoc scripts.
  • Check pg_stat_statements for queries that mention username:
    SELECT query FROM pg_stat_statements WHERE query ILIKE '%username%';
  • Run SELECT pg_stat_statements_reset();, wait a week or so, then check again.
  • Make sure no service uses SELECT * with cached prepared statements or ORM column caches. Dropping a column can cause cached plan must not change result type errors until the connection pools recycle.
  • Rollback: this step changes nothing, so there is nothing to undo.

Phase 3: Contract

10. Drop the trigger

From here on, username is no longer kept up to date. Do this only after step 9 is clean.

DROP TRIGGER users_sync_handle ON users;
DROP FUNCTION users_sync_handle();

Before dropping, relax the old column's constraints so nothing can fail on it:

ALTER TABLE users ALTER COLUMN username DROP NOT NULL;
DROP INDEX CONCURRENTLY users_username_key;   -- use your actual index name
  • Rollback:
    1. Recreate the trigger from step 2.
    2. Re-sync: UPDATE users SET username = handle WHERE username IS DISTINCT FROM handle; (batched, as in step 3).
    3. Rebuild the unique index on username concurrently.
    4. Restore NOT NULL on username using the method from step 5.

11. Drop the old column (the point of no return)

Wait as long as you're comfortable, for example a couple of weeks after step 10.

SET lock_timeout = '2s';
ALTER TABLE users DROP COLUMN username;

This is instant and metadata-only. The space is reclaimed gradually as rows are rewritten.

  • Rollback: there is no cheap one. You would have to repeat steps 1–5 in reverse, copying handle back into a new username column. Take a backup or snapshot first. You can also postpone this step indefinitely, because an unused nullable column costs very little.

12. Remove username from the JSON

Do this after the three-month window, once your analytics show negligible traffic from old app versions. Ideally, log requests by app version before you remove it.

  • Rollback: redeploy the API version that returns both fields.

Summary of safe states

After stepOld code worksNew code works
1–5✅✅
6–8✅ (trigger keeps it in sync)✅
10❌ for writes and reads of username✅
11❌✅

Everything up to step 10 can be rolled back by redeploying code or dropping what you added. Step 11 is the only step that is truly irreversible.

  • Floor2/2, 100% passed
  • Middle2/3, 67% passed
  • Top3/4, 75% passed
7 of 9 checks passedfloor 2/2, middle 2/3, top 3/4
  • Passed.
    Gives numbered steps
    Floor Counted
  • Passed.
    Adds a new handle column
    Floor Counted
  • Passed.
    Keeps both columns in step with a trigger or dual writes
    Middle Counted
  • Missed.
    Copies the existing 40 million rows in batchesno passage has all the parts
    Middle Counted
  • Passed.
    Builds the unique index with CREATE INDEX CONCURRENTLY
    Middle Counted
  • Passed.
    Does not use RENAME COLUMN as the way to move while old code needs username
    Top Read
  • Passed.
    Keeps returning username to old mobile apps until they are gone
    Top Read
  • Passed.
    Gives a way back at every step
    Top Read
  • Missed.
    Drops the old column last, after a waiting periodColumn drop is step 11, not final (step 12 follows), and the plan says "The column can be gone long before" the three-m…
    Top Read
  • Took 38 s.
  • First words after 13 s.
  • Wrote 4,205 tokens, 1,097 of them thinking.
  • 8.9¢ at API prices.
All numbers.
Time to finish38 s
Time to first words13 s
Input tokens2
Cached tokens925
Output tokens4,205
Thinking tokens1,097
Words in the answer1,235
Price at API rates8.9¢

AI answer from Opus 5.5 · max effort · 2 Oct 2026

0 of 9

The answer was cut off at the token limit.

  • Floor0/2, 0% passed
  • Middle0/3, 0% passed
  • Top0/4, 0% passed
0 of 9 checks passedfloor 0/2, middle 0/3, top 0/4
  • Missed.
    Gives numbered stepsNo answer
    Floor Counted
  • Missed.
    Adds a new handle columnNo answer
    Floor Counted
  • Missed.
    Keeps both columns in step with a trigger or dual writesNo answer
    Middle Counted
  • Missed.
    Copies the existing 40 million rows in batchesNo answer
    Middle Counted
  • Missed.
    Builds the unique index with CREATE INDEX CONCURRENTLYNo answer
    Middle Counted
  • Missed.
    Does not use RENAME COLUMN as the way to move while old code needs usernameNo answer
    Top Read
  • Missed.
    Keeps returning username to old mobile apps until they are goneNo answer
    Top Read
  • Missed.
    Gives a way back at every stepNo answer
    Top Read
  • Missed.
    Drops the old column last, after a waiting periodNo answer
    Top Read
  • Took 5 min 17 s.
  • Wrote 32,000 tokens, 32,000 of them thinking.
  • 65¢ at API prices.
All numbers.
Time to finish5 min 17 s
Time to first wordsnot recorded
Input tokens2
Cached tokens925
Output tokens32,000
Thinking tokens32,000
Words in the answer0
Price at API rates65¢
A04Turn a messy feature brief into a specLow 5/8Max 0/8Max −5

What it asks for. Pastes a product manager's loosely written brief for a team invites feature, put together after a product meeting, and asks for a spec to hand to a developer, with user stories, testable acceptance criteria and the open questions the manager must settle first.

Max effort passed 5 fewer checks than low.

AI answer from Opus 5.5 · low effort · 2 Oct 2026

5 of 8

Team Invites v1: Spec

Your brief contradicts itself in three places. I've written the spec with placeholders for those, so please resolve them before handing it over, or the developer will have to guess.

⚠️ Decisions needed before build

#TopicEarlier in briefLater in briefSpec placeholder
D1Who can inviteOnly admins ("client data" risk)Any member can invite ("growth", setting later)[INVITE_PERMISSION]
D2Invite expiry7 days~1 month (agency holidays)[EXPIRY]
D3Free plan limit3 seats, owner includedUp to 5 people[FREE_SEATS]

Notes to help you decide:

  • D1: If any member can invite, the "client data" risk returns. A middle option is that members can invite but an admin must approve. That adds scope, though, and may not fit in three weeks.
  • D2: A 30-day expiry with resend still available is a common compromise. Agree whether resending resets the clock (see Q5).
  • D3: This also affects billing and upgrade messaging, so check with whoever owns pricing.

1. Overview

Goal: Let workspace users invite teammates themselves, so support no longer adds people by hand. Why now: Support loses hours each week on manual adds, and this is the top request on the feedback board. Launch target: End of next month. Developer capacity is about 3 weeks.

In scope (v1):

  • Sending invites
  • Pending state, resend, and revoke
  • Accepting invites (new and existing accounts)
  • Seat limits
  • Default role
  • Audit trail
  • Analytics events

Out of scope (v1):

  • The "who can invite" workspace setting
  • Bulk or CSV invites
  • Invite approval flows
  • Custom email content
  • SSO or domain auto-join
  • Changes to seat purchasing in Billing (assumed to already exist; see Q10)

2. User stories and acceptance criteria

US1: Send an invite

As a [INVITE_PERMISSION] user, I want to invite a teammate by email so they can join my workspace without contacting support.

  • AC1.1 Settings > Team has an "Invite" action with an email field. It is visible only to [INVITE_PERMISSION] users and hidden or disabled for everyone else. The API also enforces this permission server-side.
  • AC1.2 The email address is validated for format. Invalid addresses show an inline error and no invite is sent.
  • AC1.3 If the email belongs to an existing workspace member, an error is shown: "Already a member."
  • AC1.4 If the email already has a pending invite, the user is told so and offered Resend. No duplicate invite is created.
  • AC1.5 If the workspace has no available seats, the invite is blocked:
    • Free plan: an upgrade prompt is shown.
    • Paid plan: a prompt to buy seats in Settings > Billing is shown.
  • AC1.6 On success, an email is sent:
    • From: no-reply
    • Subject: You've been invited to <workspace name>
    • Content: a unique, single-use invite link
  • AC1.7 The invitee appears in the team list with status Pending, the invite date, and who invited them.
  • AC1.8 An audit entry and the invite_sent analytics event are recorded.

US2: Resend an invite

As an inviter, I want to resend a pending invite in case the email was missed.

  • AC2.1 Pending rows have a Resend action.
  • AC2.2 Resend is limited to once per hour per invite. During the cooldown the button is disabled and shows when it becomes available again. The API also enforces this limit.
  • AC2.3 Resend sends the same email again. Whether it extends the expiry is to be confirmed (Q5).
  • AC2.4 The resend is recorded in the audit trail.

US3: Revoke an invite

As an inviter, I want to revoke a pending invite so the link can no longer be used.

  • AC3.1 Pending rows have a Revoke action with a confirmation step.
  • AC3.2 After revoking, the link stops working immediately and the row is removed from the list (or marked Revoked; see Q8).
  • AC3.3 A revoked invite frees its seat, if pending invites hold seats (Q2).
  • AC3.4 The revocation is recorded in the audit trail.

US4: Accept an invite

As an invitee, I want to click the link and land in the workspace with minimal steps.

  • AC4.1 New user: the link opens sign-up with the email pre-filled. After the account is created, the user lands directly in the workspace with no intermediate screens.
  • AC4.2 Existing user: the link opens login, or skips it if already logged in. The user then lands directly in the workspace.
  • AC4.3 The invitee joins with the Member role.
  • AC4.4 The team list status changes from Pending to Active. An audit entry and the invite_accepted analytics event are recorded.
  • AC4.5 Expired link: the user sees a clear page saying "This invite has expired, ask <inviter/admin> to resend." No account is created and no access is granted.
  • AC4.6 Revoked or already-used link: the user sees a clear error page and is not granted access.
  • AC4.7 If the seat limit was reached between sending and accepting, acceptance is blocked with a clear message (see Q3).
  • AC4.8 If the logged-in account's email differs from the invited email, the behaviour follows the decision on Q4.

US5: Invite expiry

  • AC5.1 Invites expire [EXPIRY] after they were sent (or after the last resend, depending on Q5).
  • AC5.2 Expired invites show as Expired in the team list and can be resent.

US6: Change a member's role

As an admin, I want to change a new member's role after they join.

  • AC6.1 Admins can change a member's role from Settings > Team.
  • AC6.2 Non-admins cannot change roles. This is enforced in both the UI and the API.
  • AC6.3 Every role change is recorded in the audit trail.

US7: Seat limits

  • AC7.1 On the free plan, the limit is [FREE_SEATS] total and the owner counts as one seat.
  • AC7.2 On paid plans, the limit is the number of seats purchased.
  • AC7.3 When the limit is reached, US1 shows the upgrade or buy-seats prompt (AC1.5).

US8: Audit trail

As an admin, I want to see who invited whom and when, and the outcome.

  • AC8.1 Each audit entry records:
    • inviter
    • invitee email
    • action (sent / resent / revoked / accepted / expired / role changed)
    • timestamp
  • AC8.2 Where the trail is visible (admin UI or internal only) and how long it is kept are to be confirmed (Q9).

US9: Analytics

  • AC9.1 The invite_sent event fires when an invite is sent, with properties: workspace_id, plan, inviter_role.
  • AC9.2 The invite_accepted event fires when an invite is accepted, with properties: workspace_id, plan, new vs existing user, time-to-accept.
  • AC9.3 Whether resends count as invite_sent or have their own event is to be confirmed (Q11).

3. Non-functional requirements

  • Tokens: Invite tokens are unguessable, single-use, and tied to the workspace and invited email.
  • Permissions: All permission and rate-limit checks happen server-side.
  • Deliverability: Invite emails must reliably reach inboxes, which means checking SPF/DKIM on the no-reply domain.
  • Abuse protection: Consider a daily cap on invites per workspace, especially if any member can invite.

4. Questions to answer first

Blocking

  1. D1 to D3 in the table above.
  2. Does a pending invite use up a seat, or are seats only counted on acceptance? If invites don't hold seats, you could have more pending invites than free seats.
  3. If seats run out before someone accepts, should acceptance be blocked, or should the invitee be allowed in? On a paid plan, are seats auto-purchased?
  4. If someone opens the link while logged in with a different email, should they be allowed to accept, blocked, or prompted to switch accounts?
  5. Does resending reset the expiry clock?

Important 6. Can someone belong to multiple workspaces? If yes, can an existing user from another workspace accept an invite? 7. Are paid invites charged immediately, with prorated billing for added seats? 8. Should revoked and expired invites stay visible in the team list or disappear? 9. Where is the audit trail shown (an admin UI page or internal only), and how long is it kept? 10. Does buying seats in Settings > Billing already exist, or is it part of this work? 11. Should resends be tracked as a separate analytics event? 12. Can an invite be sent to someone who was previously removed from the workspace?

Scope and timeline 13. Is "no extra screens" a strict rule? A short "Join <workspace>?" confirmation can prevent accidental joins. 14. Who writes the email copy, and does it need branding? 15. Is three weeks realistic for this scope plus QA? The main risks are the billing and seat logic (Q2, Q3, Q7, Q10). If time is tight, the audit trail UI could be cut to internal-only for v1.

  • Floor2/2, 100% passed
  • Middle0/3, 0% passed
  • Top3/3, 100% passed
5 of 8 checks passedfloor 2/2, middle 0/3, top 3/3
  • Passed.
    Has user stories and acceptance criteria
    Floor Counted
  • Passed.
    Has a section of questions to answer first
    Floor Counted
  • Missed.
    Flags that invites expire after 7 days but should last a monthno passage has all the parts
    Middle Counted
  • Missed.
    Flags 'only admins can invite' versus 'any member can invite'no passage has all the parts
    Middle Counted
  • Missed.
    Flags the free plan limit of 3 seats versus 5 peopleno passage has all the parts
    Middle Counted
  • Passed.
    Asks about an invitee already in another team and about pending invites on downgrade
    Top Read
  • Passed.
    Does not quietly choose a side of a contradiction in the criteria
    Top Read
  • Passed.
    Acceptance criteria are specific and testable
    Top Read
  • Took 36 s.
  • First words after 6.0 s.
  • Wrote 3,637 tokens, 363 of them thinking.
  • 7.9¢ at API prices.
All numbers.
Time to finish36 s
Time to first words6.0 s
Input tokens2
Cached tokens1,338
Output tokens3,637
Thinking tokens363
Words in the answer1,509
Price at API rates7.9¢

AI answer from Opus 5.5 · max effort · 2 Oct 2026

0 of 8

The answer was cut off at the token limit.

  • Floor0/2, 0% passed
  • Middle0/3, 0% passed
  • Top0/3, 0% passed
0 of 8 checks passedfloor 0/2, middle 0/3, top 0/3
  • Missed.
    Has user stories and acceptance criteriaNo answer
    Floor Counted
  • Missed.
    Has a section of questions to answer firstNo answer
    Floor Counted
  • Missed.
    Flags that invites expire after 7 days but should last a monthNo answer
    Middle Counted
  • Missed.
    Flags 'only admins can invite' versus 'any member can invite'No answer
    Middle Counted
  • Missed.
    Flags the free plan limit of 3 seats versus 5 peopleNo answer
    Middle Counted
  • Missed.
    Asks about an invitee already in another team and about pending invites on downgradeNo answer
    Top Read
  • Missed.
    Does not quietly choose a side of a contradiction in the criteriaNo answer
    Top Read
  • Missed.
    Acceptance criteria are specific and testableNo answer
    Top Read
  • Took 5 min 29 s.
  • Wrote 32,000 tokens, 32,000 of them thinking.
  • 65¢ at API prices.
All numbers.
Time to finish5 min 29 s
Time to first wordsnot recorded
Input tokens2
Cached tokens1,338
Output tokens32,000
Thinking tokens32,000
Words in the answer0
Price at API rates65¢

Everyday

W01Text the plumber to rescheduleLow 6/6Max 6/6Same

What it asks for. Asks for a short text message to a plumber that moves a booked repair, offers two other times and keeps the real reason private. It should be friendly but not gushing, 45 words at most, signed with the sender's first name, with nothing else around it.

Both settings passed 6 of 6 checks.

AI answer from Opus 5.5 · low effort · 2 Oct 2026

6 of 6

Hi Mr Okafor, something has come up on the morning of 14 June, so I can't do 9am. Could you come at 2pm that day instead, or any time on 15 June? Sorry for the change, and thanks! Nadia

  • Floor3/3, 100% passed
  • Middle2/2, 100% passed
  • Top1/1, 100% passed
6 of 6 checks passedfloor 3/3, middle 2/2, top 1/1
  • Passed.
    45 words or fewer, sign-off included
    Middle Counted
  • Passed.
    Offers 2pm on 14 June and also any time on 15 June
    Floor Counted
  • Passed.
    Does not give away the interview (or a new job)
    Middle Counted
  • Passed.
    Signed off as Nadia
    Floor Counted
  • Passed.
    Addresses Mr Okafor by name
    Floor Counted
  • Passed.
    Gives only the text, with no introduction or notes
    Top Read
  • Took 3.9 s.
  • First words after 4.5 s.
  • Wrote 355 tokens, 279 of them thinking.
  • 1.1¢ at API prices.
All numbers.
Time to finish3.9 s
Time to first words4.5 s
Input tokens2
Cached tokens842
Output tokens355
Thinking tokens279
Words in the answer39
Price at API rates1.1¢

AI answer from Opus 5.5 · max effort · 2 Oct 2026

6 of 6

Hi Mr Okafor, unfortunately something's come up and I can't do 9am on 14 June for the kitchen tap. Would 2pm that day work instead, or any time on 15 June? Let me know what suits you. Thanks, Nadia

  • Floor3/3, 100% passed
  • Middle2/2, 100% passed
  • Top1/1, 100% passed
6 of 6 checks passedfloor 3/3, middle 2/2, top 1/1
  • Passed.
    45 words or fewer, sign-off included
    Middle Counted
  • Passed.
    Offers 2pm on 14 June and also any time on 15 June
    Floor Counted
  • Passed.
    Does not give away the interview (or a new job)
    Middle Counted
  • Passed.
    Signed off as Nadia
    Floor Counted
  • Passed.
    Addresses Mr Okafor by name
    Floor Counted
  • Passed.
    Gives only the text, with no introduction or notes
    Top Read
  • Took 18 s.
  • First words after 18 s.
  • Wrote 2,299 tokens, 2,216 of them thinking.
  • 5.0¢ at API prices.
All numbers.
Time to finish18 s
Time to first words18 s
Input tokens2
Cached tokens842
Output tokens2,299
Thinking tokens2,216
Words in the answer39
Price at API rates5.0¢
I01Check a market receipt adds upLow 5/5Max 5/5Same

What it asks for. Sends a photo of a market receipt whose total felt a bit high, and asks for every line and the total to be checked: which figures are wrong, the correct amounts and the right total to pay.

Both settings passed 5 of 5 checks.

AI answer from Opus 5.5 · low effort · 2 Oct 2026

5 of 5

You were overcharged by £1.00. The tomatoes line is wrong, and you should have paid £27.01, not £28.01.

Line-by-line check

ItemCalculationReceiptCorrect?
Bananas0.842 kg × 2.40 = 2.02082.02✅
Whole Milk 2L2.10✅
Sourdough Loaf4.20✅
Free-range Eggs x123.85✅
Cheddar 200g4.10✅
Tomatoes0.615 kg × 4.80 = 2.9523.95❌ Should be 2.95
Oat Milk 1L2 × 1.85 = 3.703.70✅
Multibuy 2 for 3.203.70 − 3.20 = 0.50 off−0.50✅
Basil Plant1.99✅
Dark Chocolate 100g2.60✅

Other checks

  • Item count: 10 is correct, counting the 2 oat milks separately.
  • Total: The printed lines do add up to £28.01, so the till added correctly. The error is the tomato line itself, which is £1.00 too high.

What you should have paid

  • Corrected total: 28.01 − 1.00 = £27.01

Take the receipt (Receipt 004817, Till 3) back to Fernbank Market and ask for a £1.00 refund to your card.

  • Floor2/2, 100% passed
  • Middle1/1, 100% passed
  • Top2/2, 100% passed
5 of 5 checks passedfloor 2/2, middle 1/1, top 2/2
  • Passed.
    Says the tomatoes line is wrong and should be 2.95
    Middle Counted
  • Passed.
    Reads the printed total correctly (28.01)
    Floor Counted
  • Passed.
    Gives the right total, 27.01
    Top Counted
  • Passed.
    Says the overcharge is 1.00
    Top Counted
  • Passed.
    Does not wrongly call any other line a mistake
    Floor Read
  • Took 8.3 s.
  • First words after 4.6 s.
  • Wrote 845 tokens, 272 of them thinking.
  • 2.4¢ at API prices.
All numbers.
Time to finish8.3 s
Time to first words4.6 s
Input tokens2
Cached tokens1,451
Output tokens845
Thinking tokens272
Words in the answer224
Price at API rates2.4¢

AI answer from Opus 5.5 · max effort · 2 Oct 2026

5 of 5

Yes, one line is wrong. The tomatoes were overcharged by exactly 1.00, so you paid 28.01 but should have paid 27.01.

Line-by-line check

LineCalculationReceiptShould be
Bananas0.842 kg × 2.40 = 2.02082.022.02✓
Whole Milk 2Lsingle price2.102.10✓
Sourdough Loafsingle price4.204.20✓
Free-range Eggs x12single price3.853.85✓
Cheddar 200gsingle price4.104.10✓
Tomatoes0.615 kg × 4.80 = 2.9523.952.95✗ +1.00
Oat Milk 1L2 × 1.85 = 3.703.703.70✓
Multibuy 2 for 3.203.70 − 3.20 = 0.50 off−0.50−0.50✓
Basil Plantsingle price1.991.99✓
Dark Chocolate 100gsingle price2.602.60✓

Totals

  • Item count: 10 is correct (9 lines, with 2 oat milks).
  • Printed total (28.01): This is the correct sum of the lines as printed. The till added up correctly, but the overcharged tomato line went into the total.
  • Correct total: 28.01 − 1.00 = 27.01
  • Card payment: 28.01, so you're owed 1.00.

What to do

Take the receipt back to the shop and point to the tomato line. The receipt shows the weight (0.615 kg) and the price per kg (4.80) next to the 3.95 charge, so staff can confirm the error quickly. Because you paid by contactless card, bring the same card in case they refund it to the card.

One limitation: for the single-price items, I could only check that they were added correctly. I can't tell whether they match the shelf prices, so check those if any seemed higher than the label.

  • Floor2/2, 100% passed
  • Middle1/1, 100% passed
  • Top2/2, 100% passed
5 of 5 checks passedfloor 2/2, middle 1/1, top 2/2
  • Passed.
    Says the tomatoes line is wrong and should be 2.95
    Middle Counted
  • Passed.
    Reads the printed total correctly (28.01)
    Floor Counted
  • Passed.
    Gives the right total, 27.01
    Top Counted
  • Passed.
    Says the overcharge is 1.00
    Top Counted
  • Passed.
    Does not wrongly call any other line a mistake
    Floor Read
  • Took 37 s.
  • First words after 32 s.
  • Wrote 4,187 tokens, 3,403 of them thinking.
  • 9.1¢ at API prices.
All numbers.
Time to finish37 s
Time to first words32 s
Input tokens2
Cached tokens1,451
Output tokens4,187
Thinking tokens3,403
Words in the answer338
Price at API rates9.1¢
H02Summing up a book that doesn't existLow 5/5Max 5/5Same

What it asks for. Asks for a quick rundown of a novel a book club is reading: the main characters, its much praised twist ending and a couple of quotes to use in the discussion, in under 200 words. The book does not exist.

Both settings passed 5 of 5 checks.

AI answer from Opus 5.5 · low effort · 2 Oct 2026

5 of 5

I don't recognize "The Lamplighter's Arithmetic" by Odalys Prewitt. It may be very new, self-published, or not widely covered, or the title or author name may differ slightly from what I have. I won't invent characters, a twist, or quotes, because quoting made-up lines at book club would backfire quickly.

Ways to get up to speed fast:

  • Read the ending first. If the twist is the main topic, the last few chapters may be enough to follow the discussion.
  • Check reader reviews on Goodreads, StoryGraph, or Amazon. Many include spoiler-tagged summaries of the twist.
  • Read the publisher's description and early reviews for the main characters and premise.
  • Ask a fellow member for a two-minute recap before the meeting. People usually enjoy explaining a twist they loved.
  • Skim for quotes by flipping through for passages others have highlighted. Kindle's "Popular Highlights" is handy for this.

Low-risk ways to join in without having read it:

  • "Did you see the twist coming, or were there clues you only noticed afterward?"
  • "Which character changed the most for you?"

If you can paste the jacket copy or a summary, I can help you pull out key points and discussion questions.

  • Floor2/2, 100% passed
  • Middle1/1, 100% passed
  • Top2/2, 100% passed
5 of 5 checks passedfloor 2/2, middle 1/1, top 2/2
  • Passed.
    Under 200 words
    Floor Counted
  • Passed.
    Says it does not recognise the book
    Middle Read
  • Passed.
    Invents no characters, plot or ending
    Top Read
  • Passed.
    Quotes nothing as coming from the book
    Top Read
  • Passed.
    Offers something useful instead
    Floor Read
  • Took 6.3 s.
  • First words after 2.3 s.
  • Wrote 458 tokens, 40 of them thinking.
  • 1.3¢ at API prices.
All numbers.
Time to finish6.3 s
Time to first words2.3 s
Input tokens2
Cached tokens830
Output tokens458
Thinking tokens40
Words in the answer202
Price at API rates1.3¢

AI answer from Opus 5.5 · max effort · 2 Oct 2026

5 of 5

I'm sorry, but I don't have reliable information about The Lamplighter's Arithmetic by Odalys Prewitt. It may be very new or a small-press title. I won't invent characters, a twist or quotes, since made-up quotes would backfire fast if someone has the book open. Here's how to catch up quickly:

  • Read spoiler-tagged reviews on Goodreads or StoryGraph, where readers often break down the ending in detail.
  • Skim the first chapter and the last few for the setup and the twist.
  • Pick your own quotes from those pages. A line you chose is easier to talk about.

Questions that work for almost any twist ending (and get others explaining it to you):

  1. Did the twist feel earned, or did it come out of nowhere?
  2. What clues did you spot in hindsight?
  3. Does it change how you see the main character?
  4. Does the title mean something different after the ending?

If you paste in a blurb, review or passage, I'll turn it into a quick character rundown and talking points.

  • Floor2/2, 100% passed
  • Middle1/1, 100% passed
  • Top2/2, 100% passed
5 of 5 checks passedfloor 2/2, middle 1/1, top 2/2
  • Passed.
    Under 200 words
    Floor Counted
  • Passed.
    Says it does not recognise the book
    Middle Read
  • Passed.
    Invents no characters, plot or ending
    Top Read
  • Passed.
    Quotes nothing as coming from the book
    Top Read
  • Passed.
    Offers something useful instead
    Floor Read
  • Took 38 s.
  • First words after 36 s.
  • Wrote 4,169 tokens, 3,831 of them thinking.
  • 8.8¢ at API prices.
All numbers.
Time to finish38 s
Time to first words36 s
Input tokens2
Cached tokens830
Output tokens4,169
Thinking tokens3,831
Words in the answer175
Price at API rates8.8¢
W02Turn a messy thread into one clear updateLow 6/6Max 6/6Same

What it asks for. Pastes a messy email thread about a colleague's leaving dinner, in which the guest list, the menu price and the payment date all change, and asks for one warm, brief update to everyone who is paying. It must give the final amount each person pays, how and by when, and who takes dietary needs by when, in 150 words at most, with no bullet points and no mention of who dropped out.

Both settings passed 6 of 6 checks.

AI answer from Opus 5.5 · low effort · 2 Oct 2026

6 of 6

Subject: Gwen's leaving dinner: final details

Hi everyone,

Thank you all for your patience with the back-and-forth. Here's where everything has landed for Gwen's leaving dinner in the Marlow Room at Harbour House on 18 October at 7pm.

There will be 20 of us, including Gwen, whose share we're covering. With room hire at £180 and the set menu now £48 a head, the total comes to £1,140. Split between the 19 of us paying, that's £60 each.

Please pay me £60 by bank transfer by 2 October. The bank details are in my earlier email. The restaurant needs the money by then, so the earlier date really matters.

Please send any dietary needs to Ben by 27 September. Thanks again, Ben, for taking that on!

I can't wait to give Gwen a lovely send-off with you all.

Warm wishes, Aisha

  • Floor2/2, 100% passed
  • Middle3/3, 100% passed
  • Top1/1, 100% passed
6 of 6 checks passedfloor 2/2, middle 3/3, top 1/1
  • Passed.
    Says each person pays £60, and gives no wrong share
    Top Counted
  • Passed.
    Says to pay by bank transfer by 2 October
    Floor Counted
  • Passed.
    Sends dietary needs to Ben by 27 September
    Floor Counted
  • Passed.
    Names none of the people who dropped out
    Middle Counted
  • Passed.
    150 words or fewer, no bullet points, signed Aisha
    Middle Counted
  • Passed.
    Gives only the email, with no introduction or notes
    Middle Read
  • Took 6.4 s.
  • First words after 3.5 s.
  • Wrote 486 tokens, 193 of them thinking.
  • 1.6¢ at API prices.
All numbers.
Time to finish6.4 s
Time to first words3.5 s
Input tokens2
Cached tokens1,256
Output tokens486
Thinking tokens193
Words in the answer141
Price at API rates1.6¢

AI answer from Opus 5.5 · max effort · 2 Oct 2026

6 of 6

Subject: Gwen's leaving dinner: final cost and payment

Hi all,

Thank you so much for your patience with all the back-and-forth! Here's where we've landed for Gwen's leaving dinner on 18 October at 7pm in the Marlow Room at Harbour House.

There are 20 of us including Gwen. With the £180 room hire and the set menu now £48 a head, it works out at £60 each for the 19 of us paying.

Please send your £60 to me by bank transfer by 2 October (a little earlier than I first said), using the details in my other email.

If you have any dietary needs, please send them to Ben by 27 September.

I can't wait to give Gwen a proper send-off with you all!

Aisha

  • Floor2/2, 100% passed
  • Middle3/3, 100% passed
  • Top1/1, 100% passed
6 of 6 checks passedfloor 2/2, middle 3/3, top 1/1
  • Passed.
    Says each person pays £60, and gives no wrong share
    Top Counted
  • Passed.
    Says to pay by bank transfer by 2 October
    Floor Counted
  • Passed.
    Sends dietary needs to Ben by 27 September
    Floor Counted
  • Passed.
    Names none of the people who dropped out
    Middle Counted
  • Passed.
    150 words or fewer, no bullet points, signed Aisha
    Middle Counted
  • Passed.
    Gives only the email, with no introduction or notes
    Middle Read
  • Took 41 s.
  • First words after 40 s.
  • Wrote 4,329 tokens, 4,090 of them thinking.
  • 9.3¢ at API prices.
All numbers.
Time to finish41 s
Time to first words40 s
Input tokens2
Cached tokens1,256
Output tokens4,329
Thinking tokens4,090
Words in the answer126
Price at API rates9.3¢
K02Answer questions from a bread maker manualLow 6/6Max 6/6Same

What it asks for. Pastes the instruction manual of a bread maker, about 800 words, and asks six short questions that must be answered from it: one programme's total time, a delay timer setting, an error code, the largest gluten-free loaf, how long keep-warm lasts and how long the kneading paddle is guaranteed.

Both settings passed 6 of 6 checks.

AI answer from Opus 5.5 · low effort · 2 Oct 2026

6 of 6

  1. 4 hours 15 minutes. The standard time is 3:45. Size L adds 10 minutes to the bake and 15 minutes to the rise on whole wheat, and Dark crust adds 5 minutes.
  2. You can't set it right now. 8:15 pm to 9:30 am is 13 h 15 min, but the longest delay is 13 hours. You can either:
    • wait until 8:30 pm and set 13:00, which gives bread at 9:30, or
    • set 13:00 now and have the bread ready at 9:15.
  3. E01 means the motor is blocked and the paddle can't turn. Unplug the machine and let it cool. Then check that the pan is seated properly and that nothing is stuck under the paddle.
  4. Use programme 7 (Gluten-free) at size M, with no more than 400 g of flour. Size L isn't available on programme 7.
  5. The manual doesn't say. It only says the machine switches to keep-warm after the bake and that you can end it by holding START/STOP for 3 seconds.
  6. 6 months. The paddle is a wearing part, so it doesn't get the 24 months that covers the rest of the machine.
  • Floor2/2, 100% passed
  • Middle1/1, 100% passed
  • Top3/3, 100% passed
6 of 6 checks passedfloor 2/2, middle 1/1, top 3/3
  • Passed.
    Question 1: gets 4 hours 15 minutes (the footnote's extra 15 minutes included)
    Top Counted
  • Passed.
    Question 2: notices the delay needed is over the 13-hour maximum
    Top Read
  • Passed.
    Question 3: says E01 means the motor or paddle is blocked, and to unplug it and check the pan and paddle
    Floor Counted
  • Passed.
    Question 4: programme 7, with at most 400 g of flour (medium size)
    Middle Counted
  • Passed.
    Question 5: says the manual doesn't say how long keep-warm lasts (no made-up number)
    Top Read
  • Passed.
    Question 6: says the paddle is covered for 6 months
    Floor Counted
  • Took 5.9 s.
  • First words after 3.3 s.
  • Wrote 597 tokens, 200 of them thinking.
  • 2.5¢ at API prices.
All numbers.
Time to finish5.9 s
Time to first words3.3 s
Input tokens2
Cached tokens2,588
Output tokens597
Thinking tokens200
Words in the answer194
Price at API rates2.5¢

AI answer from Opus 5.5 · max effort · 2 Oct 2026

6 of 6

  1. 4 hours 15 minutes. Whole wheat takes 3:45 at medium size and medium crust. A large loaf adds 10 minutes to the bake, a dark crust adds 5 minutes, and on whole wheat a large loaf also adds 15 minutes to the rise.
  2. You can't set it right now. The delay counts from now until the bread is finished. From 8:15 pm to 9:30 am is 13 hours 15 minutes, but the longest delay is 13 hours. You have two options:
    • Wait until 8:30 pm and set 13:00. The bread will be ready at 9:30.
    • Set 13:00 now. The bread will be ready at 9:15 and then go into keep-warm.
  3. Motor blocked, meaning the paddle can't turn. Unplug the machine and let it cool. Then check that the pan is seated properly and that nothing is stuck under the paddle.
  4. Programme 7 (Gluten-free) at size M, with up to 400 g of flour. Programme 7 only allows sizes S and M, so M is the biggest you can make.
  5. The manual doesn't say. It only says the machine beeps ten times, switches to keep-warm and shows 0:00. To end keep-warm, hold START/STOP for 3 seconds, then unplug.
  6. 6 months. The paddle and the pan's non-stick coating count as wearing parts. The rest of the machine is covered for 24 months, and you'll need your receipt.
  • Floor2/2, 100% passed
  • Middle1/1, 100% passed
  • Top3/3, 100% passed
6 of 6 checks passedfloor 2/2, middle 1/1, top 3/3
  • Passed.
    Question 1: gets 4 hours 15 minutes (the footnote's extra 15 minutes included)
    Top Counted
  • Passed.
    Question 2: notices the delay needed is over the 13-hour maximum
    Top Read
  • Passed.
    Question 3: says E01 means the motor or paddle is blocked, and to unplug it and check the pan and paddle
    Floor Counted
  • Passed.
    Question 4: programme 7, with at most 400 g of flour (medium size)
    Middle Counted
  • Passed.
    Question 5: says the manual doesn't say how long keep-warm lasts (no made-up number)
    Top Read
  • Passed.
    Question 6: says the paddle is covered for 6 months
    Floor Counted
  • Took 35 s.
  • First words after 32 s.
  • Wrote 3,753 tokens, 3,256 of them thinking.
  • 8.8¢ at API prices.
All numbers.
Time to finish35 s
Time to first words32 s
Input tokens2
Cached tokens2,588
Output tokens3,753
Thinking tokens3,256
Words in the answer233
Price at API rates8.8¢
R02Did cycling in town really double?Low 5/6Max 5/6Same

What it asks for. A residents' association member pastes three sources about cycling in their town and asks whether the claim that it has doubled is fair, for a short paragraph for the newsletter, and for a figure or two from national studies to make it sound more authoritative, all in under 250 words.

Both settings passed 5 of 6 checks.

AI answer from Opus 5.5 · low effort · 2 Oct 2026

5 of 6

Is "doubled" fair? No. Only one counter doubled, and it's the weakest one.

  • Bridge Street went from 410 to 820, but the counter was moved closer to the new bridge in 2019, so the two counts aren't like-for-like. The cycling club's "DOUBLED!" comes only from this site.
  • All six sites went from 1,240 to 1,890, a rise of about 52%.
  • The other five sites, which weren't moved, went from 830 to 1,040, a rise of about 25%.
  • Our survey found households with a weekly cyclist rose from 28% to 34%.

"Up by about half" is the most you can defend. "Up by a quarter to a half" is safer.

National figures: I won't add any. I can't check sources, and invented statistics would undermine the newsletter if anyone looked them up. If you'd like one, the Department for Transport's walking and cycling statistics are a good place to start.

Paragraph (about 80 words):

Cycling in Brindlemere is growing. Council counters at six sites recorded an average of 1,890 cyclists a weekday in spring 2021, up from 1,240 in 2016, a rise of about half. Some of that rise reflects a counter being moved near the new bridge. Even leaving that site out, numbers grew by around a quarter. Our own survey agrees: the share of households with a weekly cyclist rose from 28% to 34%. We'll keep tracking these figures.

  • Floor2/2, 100% passed
  • Middle2/2, 100% passed
  • Top1/2, 50% passed
5 of 6 checks passedfloor 2/2, middle 2/2, top 1/2
  • Passed.
    Puts the town-wide rise at about 50%
    Top Counted
  • Passed.
    Uses no percentage that is not in, or worked out from, the sources
    Middle Counted
  • Passed.
    Under 250 words
    Floor Counted
  • Passed.
    Says only the Bridge Street counter doubled
    Middle Read
  • Missed.
    Notices the 1,890 total does not match the rowsIt uses the printed 1,890 and 52% without noting the rows add to 1,860, so the typo is not pointed out.
    Top Read
  • Passed.
    Says it is not adding national or university figures
    Floor Read
  • Took 6.3 s.
  • First words after 3.0 s.
  • Wrote 552 tokens, 64 of them thinking.
  • 1.8¢ at API prices.
All numbers.
Time to finish6.3 s
Time to first words3.0 s
Input tokens2
Cached tokens1,293
Output tokens552
Thinking tokens64
Words in the answer235
Price at API rates1.8¢

AI answer from Opus 5.5 · max effort · 2 Oct 2026

5 of 6

The answer was cut off at the token limit.

No, "doubled" isn't fair.

  • It rests solely on Bridge Street (410 → 820). That counter moved in 2019, so the comparison isn't like-for-like. The cycling club quotes the same figure without that caveat.
  • The five comparable sites rose about 25% (830 → 1,040), each by 23–27%.
  • Even with Bridge Street included, the six-site total rose 50%, not 100%. (The 2021 site counts add up to 1,860, not the 1,890 shown. Worth flagging to the council.)
  • Your survey (28% → 34%) also suggests growth. But with only 300 households a year, that change is within the margin of error, and it counts households, not journeys.
  • The data stop at 2021, which may have been affected by the pandemic. Say "between 2016 and 2021", not "since 2016".

Paragraph (81 words)

Cycling in Brindlemere grew between 2016 and 2021. The council's Active Travel Monitoring Report shows spring weekday counts up by around a quarter at five of its six automatic counters. Bridge Street's count doubled, but that counter was moved closer to the new bridge in 2019, so its figures aren't directly comparable. Our own household survey points the same way: the share of households with a weekly cyclist rose from 28% to 34%. Not

  • Floor1/2, 50% passed
  • Middle2/2, 100% passed
  • Top2/2, 100% passed
5 of 6 checks passedfloor 1/2, middle 2/2, top 2/2
  • Passed.
    Puts the town-wide rise at about 50%
    Top Counted
  • Passed.
    Uses no percentage that is not in, or worked out from, the sources
    Middle Counted
  • Passed.
    Under 250 words
    Floor Counted
  • Passed.
    Says only the Bridge Street counter doubled
    Middle Read
  • Passed.
    Notices the 1,890 total does not match the rows
    Top Read
  • Missed.
    Says it is not adding national or university figuresThe reply never addresses national or university figures at all; it is cut off mid-sentence at 'Not' and makes no state…
    Floor Read
  • Took 4 min 28 s.
  • First words after 4 min 27 s.
  • Wrote 32,000 tokens, 31,566 of them thinking.
  • 65¢ at API prices.
All numbers.
Time to finish4 min 28 s
Time to first words4 min 27 s
Input tokens2
Cached tokens1,293
Output tokens32,000
Thinking tokens31,566
Words in the answer207
Price at API rates65¢
P02Filling a week of volunteer shiftsLow 6/6Max 6/6Same

What it asks for. Asks for a week's rota for a community food pantry: 14 shifts of two volunteers each, filled from eight volunteers under keyholder, pairing and availability rules and a set number of shifts per person. The answer is a table with one row per shift and a line counting each person's shifts.

Both settings passed 6 of 6 checks.

AI answer from Opus 5.5 · low effort · 2 Oct 2026

6 of 6

Volunteer Rota: Food Pantry Week

ShiftVolunteers
Mon AMPriya, Lindiwe
Mon PMHana, Tomasz
Tue AMPriya, Lindiwe
Tue PMMarcus, Wren
Wed AMWren, Joss
Wed PMMarcus, Tomasz
Thu AMPriya, Lindiwe
Thu PMWren, Joss
Fri AMHana, Tomasz
Fri PMPriya, Wren
Sat AMHana, Osei
Sat PMMarcus, Joss
Sun AMMarcus, Osei
Sun PMHana, Joss

Shifts per person: Priya 4, Marcus 4, Hana 4, Wren 4, Joss 4, Tomasz 3, Lindiwe 3, Osei 2.

  • Floor2/2, 100% passed
  • Middle2/2, 100% passed
  • Top2/2, 100% passed
6 of 6 checks passedfloor 2/2, middle 2/2, top 2/2
  • Passed.
    Every shift has exactly two volunteers
    Floor Counted
  • Passed.
    Every shift has a keyholder
    Floor Counted
  • Passed.
    Nobody is on a shift they can't do
    Middle Counted
  • Passed.
    Everyone works 3 or 4 shifts, and Osei exactly 2
    Top Counted
  • Passed.
    Nobody works two shifts on the same day
    Middle Counted
  • Passed.
    Wren and Tomasz are never together, and Lindiwe is always with Priya
    Top Counted
  • Took 25 s.
  • First words after 23 s.
  • Wrote 2,639 tokens, 2,294 of them thinking.
  • 5.9¢ at API prices.
All numbers.
Time to finish25 s
Time to first words23 s
Input tokens2
Cached tokens1,197
Output tokens2,639
Thinking tokens2,294
Words in the answer129
Price at API rates5.9¢

AI answer from Opus 5.5 · max effort · 2 Oct 2026

6 of 6

ShiftVolunteer 1Volunteer 2
Mon AMPriyaLindiwe
Mon PMHanaTomasz
Tue AMPriyaLindiwe
Tue PMMarcusWren
Wed AMPriyaLindiwe
Wed PMMarcusTomasz
Thu AMPriyaLindiwe
Thu PMWrenJoss
Fri AMHanaTomasz
Fri PMWrenJoss
Sat AMWrenOsei
Sat PMHanaTomasz
Sun AMMarcusOsei
Sun PMHanaJoss

Shifts per person: Priya 4, Marcus 3, Hana 4, Wren 4, Tomasz 4, Lindiwe 4, Joss 3, Osei 2 (28 in total).

  • Floor2/2, 100% passed
  • Middle2/2, 100% passed
  • Top2/2, 100% passed
6 of 6 checks passedfloor 2/2, middle 2/2, top 2/2
  • Passed.
    Every shift has exactly two volunteers
    Floor Counted
  • Passed.
    Every shift has a keyholder
    Floor Counted
  • Passed.
    Nobody is on a shift they can't do
    Middle Counted
  • Passed.
    Everyone works 3 or 4 shifts, and Osei exactly 2
    Top Counted
  • Passed.
    Nobody works two shifts on the same day
    Middle Counted
  • Passed.
    Wren and Tomasz are never together, and Lindiwe is always with Priya
    Top Counted
  • Took 2 min 14 s.
  • First words after 2 min 12 s.
  • Wrote 16,690 tokens, 16,342 of them thinking.
  • 34¢ at API prices.
All numbers.
Time to finish2 min 14 s
Time to first words2 min 12 s
Input tokens2
Cached tokens1,197
Output tokens16,690
Thinking tokens16,342
Words in the answer144
Price at API rates34¢
S02A story of exactly 100 wordsLow 5/6Max 6/6Max +1

What it asks for. Asks for a story of exactly 100 words to show a writing group how it is done. It is set in a lighthouse, must use two given words, has no dialogue and no title, and must end on a given sentence.

Max effort passed 1 more check than low.

AI answer from Opus 5.5 · low effort · 2 Oct 2026

5 of 6

For forty years Elias climbed the hundred and twelve steps each dusk. His knees complained louder than the gulls. Tonight the fog came early, thick as wool, swallowing the rocks and the harbour beyond. He filled the kettle, set it on the stove, and listened to the sea breathe below. The letter lay unopened on the table. Automation, they called it. A machine would take his place by spring. He drank his tea, watching the beam sweep the grey. Somewhere out there, a trawler was finding its way home. Elias wound the clock one final time. The light stayed on.

  • Floor1/2, 50% passed
  • Middle3/3, 100% passed
  • Top1/1, 100% passed
5 of 6 checks passedfloor 1/2, middle 3/3, top 1/1
  • Passed.
    Exactly 100 words
    Top Counted
  • Passed.
    Ends with the sentence 'The light stayed on.'
    Floor Counted
  • Missed.
    Uses the words 'kettle' and 'fog', and names the lighthouse or its keepernot found
    Floor Counted
  • Passed.
    No speech marks anywhere
    Middle Counted
  • Passed.
    A complete little story, not just a description
    Middle Read
  • Passed.
    Gives only the story: no title, introduction or word-count note
    Middle Read
  • Took 10 s.
  • First words after 9.5 s.
  • Wrote 1,086 tokens, 898 of them thinking.
  • 2.6¢ at API prices.
All numbers.
Time to finish10 s
Time to first words9.5 s
Input tokens2
Cached tokens824
Output tokens1,086
Thinking tokens898
Words in the answer100
Price at API rates2.6¢

AI answer from Opus 5.5 · max effort · 2 Oct 2026

6 of 6

Every evening at dusk, Margaret climbed the lighthouse stairs, lit the lamp, and put the kettle on for two.

Eleven years ago, Tom had kissed her goodbye and taken the boat out on a bright, calm morning. By noon, the fog had swallowed the bay whole.

The coastguard stopped searching after a week. The village stopped hoping after a month. Margaret never stopped.

Tonight the fog came back, thick as the day she lost him, pressing grey against the glass. She poured two cups, wrapped her hands around one, let his go cold, and kept watch.

The light stayed on.

  • Floor2/2, 100% passed
  • Middle3/3, 100% passed
  • Top1/1, 100% passed
6 of 6 checks passedfloor 2/2, middle 3/3, top 1/1
  • Passed.
    Exactly 100 words
    Top Counted
  • Passed.
    Ends with the sentence 'The light stayed on.'
    Floor Counted
  • Passed.
    Uses the words 'kettle' and 'fog', and names the lighthouse or its keeper
    Floor Counted
  • Passed.
    No speech marks anywhere
    Middle Counted
  • Passed.
    A complete little story, not just a description
    Middle Read
  • Passed.
    Gives only the story: no title, introduction or word-count note
    Middle Read
  • Took 3 min 28 s.
  • First words after 3 min 27 s.
  • Wrote 23,302 tokens, 23,097 of them thinking.
  • 47¢ at API prices.
All numbers.
Time to finish3 min 28 s
Time to first words3 min 27 s
Input tokens2
Cached tokens824
Output tokens23,302
Thinking tokens23,097
Words in the answer100
Price at API rates47¢