</p>

It can't tell you when the work got worse, and it can't tell you when someone did something remarkable. Both are judgments about the output, and judgment is exactly what a dashboard is built to avoid.

I learned this on both sides. At Typeface I watched a few people push past the comfortable edge of what these tools could do and come back with things no straight line of coding would have produced. One rebuilt an entire app himself, because the gap between his taste and the thing he could ship had collapsed. What set those people apart was never how many tokens they spent. The tokens were an artifact. The shift was starting from a real problem and wanting a 10x answer to it.

Then I turn to our customers and hear the same thing. A marketing leader tells me the AI content wasn't good enough. I ask how they measured quality. A pause, and then: "We could just tell."

Most AI scorecards count motion. Usage up, seats filled, headcount down, the chart is green. None of it touches whether the work is any good, which is the only question worth asking. I wrote about why the most valuable work, the breakthroughs and the judgment behind them, never shows up on a dashboard, and the one test that does measure the right thing.

Full piece is on Substack. Link in comments.