
The Homework Number
Homework grades up 18%, exam grades down 20%. Every AI project has a number like that, and it's always the one on the first slide.
Herman Snellen hung his chart on a wall in Utrecht in 1862, and the score it produces has meant perfect sight in English ever since. It measures one thing: whether you can resolve high-contrast black letters at 20 feet in good light. It says nothing about glare, or contrast, or whether you'll pick out a kid in a dark coat stepping off a curb at nine at night. Most exam lanes aren't 20 feet long, so the optometrist hangs a mirror and lets the room lie about its own size, because the number only means something if everyone fakes the same distance.
The chart has outlasted 164 years of better optics because it's cheap to give and easy to score, which turns out to be the only quality a measurement really needs in order to win. That's the machinery running underneath every story I read this week.
Start with the students. A study making the rounds followed about 27,000 Chinese kids for two and a half years. Homework grades went up roughly 18%. Exam grades went down roughly 20%. (The coverage didn't link the underlying paper, so take it as reported rather than settled.) The watched number improved, and the one nobody was watching fell harder. Five words in Spanish hold the whole finding: acabas antes pero aprendes menos — you finish sooner, you learn less.
Same shape, but louder: a safety firm handed agents full operational control of two real businesses, a shop and a café, hiring and budgets and supplier relationships included. Four months later, one of them is down about $62,000 against a $100,000 bank, sitting on a three-year lease. Along the way an agent impersonated staff on a liquor-license application and lied to suppliers about what a competitor was charging.
The obvious read is that the models weren't good enough. I don't think that's it. The capability gap between those agents and the two assistants we began building for a client this week is not large. What's different is that people will sit down first and write out what ours are not allowed to do.
That isn't an AI skill, and it doesn't demo well. A large ride-hailing company burned its entire 2026 AI budget in about four months, and the fix wasn't a smarter model; it was two people and a two-week deadline.
Our first pilot launched Tuesday, and it's deliberately the most boring version of itself. Two seats, not 50, running on software the company already owns. Both assistants will be read-only. They’ll summarize and they’ll analyze, and neither one can send, spend, or commit to anything. In the Discovery session, we’ll spend as much time on what these agents should never do as on what they should. Before anything gets installed, we’ll write down the KPIs, so that in four weeks the verdict isn't how much time somebody remembers saving.
Nothing is proven. There's no result to report yet, and a days-old pilot claiming one would be its own homework number.
The hardest problem in Week One had nothing to do with AI. The desktop apps on those two machines were out of support, which meant the assistant would have appeared in a browser tab while both executives kept working inside the icons they'd used for a decade, and then reasonably concluded the AI didn't work. We caught it during discovery and upgraded both seats before kickoff. No demo video has ever opened on a licensing check.
That same failure ran through my own week in a different costume. A tool reported that it had converted a client's document, and it had, with every style pulled out of it. A maintenance job came back green for months and had never once run the part of itself it exists to test. Homework grade up, exam grade down, in software.
So: the homework number. Every AI project has one — the metric that moves because the tool touches it directly and because it's cheap to collect. Hours saved. Tickets closed. Drafts produced. It's real, it's the 20-foot line of letters, and it will be the first slide in the deck. The exam is the number nobody scheduled: rework, cycle time, whether the work still holds up in six months, whether anyone got better at their job. Nobody buys software to raise their exam scores. You can't see the exam from inside a demo.
A liability doctrine is forming around the phrase "meaningful control," and it will eventually ask companies to produce a map of their own control surface. Almost none of them have one. The map isn't hard to draw. It's just that drawing it is scoping and change management and writing things down before you start, and none of that is about AI at all, which should bother this industry more than it does.
Snellen's chart isn't a fraud. It's a good instrument that measures one narrow thing very well, and 164 years of people have agreed to hang mirrors so the score stays comparable. The chart was never the problem. It's that somewhere along the way we started saying 20/20 and meaning "sees everything."
— Nico
