2026-09-04 · 5 min read · performance-culture

You rolled out AI. Can you tell if performance actually improved?

Usage is up and to the right. That number cannot answer the question you actually have, and the research says the feeling of speed points the wrong way.

You mandated AI. Licences went out, the training happened, adoption is up and to the right on the slide.

Now the honest question. Can you tell if anyone's actual output got better?

Most founders I ask go quiet for a second on that one. Not because they're hiding anything. Because the true answer is "I think so" and they know "I think so" isn't an answer.

The uncomfortable part is that the data is worse than "unclear"#

Two findings, and they point the same way.

Some of the output is actively negative. Around 40% of workers reported receiving AI-generated work that looked finished but had to be redone, in research from BetterUp Labs with Stanford's Social Media Lab. The researchers call it workslop. It looks like completed work at the point of handoff, so it clears the bar of "looks productive," and then it costs somebody downstream about two hours to fix. Your usage dashboard counts that as a win.

And the felt sense of speed lies in a specific direction. In a randomised controlled trial of 16 experienced developers working on repositories they knew well, they were about 19% slower when using AI tools, while estimating they had been roughly 20% faster. Small study, and the authors are careful to say it doesn't generalise to every task or every team. Take it as one solid data point rather than a law. But note what kind of point it is: those people weren't slightly off about their own speed. They were wrong about the sign.

Put those two together and you get an unpleasant shape. Usage climbs. The subjective sense of productivity climbs. Real output stays flat or dips. Every instrument on your dashboard is reading the first two.

Why you can't see it, and it isn't AI's fault#

Here's the thing I'd want you to take away even if you never talk to me again.

You're measuring adoption because adoption is measurable. The thing you actually want to know is whether judgment work got better, and that was never measurable, including in 2019, before any of this.

You didn't previously notice the gap, because nothing was forcing the question. Then you spent real money on a mandate, and now somebody is going to ask you what it bought. AI didn't create the measurement blind spot. It made it expensive.

That reframe matters, because it tells you which problem to go solve. If you believe AI broke your measurement, you'll go shopping for an AI measurement tool. If you see that measurement was already broken, you'll go looking for something else entirely.

The trap most vendors will sell you#

The pitch you'll get is some version of: mandate a target percentage of AI-assisted work, then show me who's using it and who isn't.

Two problems.

The first is what it does to your people. It turns a capability rollout into surveillance, and everybody can feel the difference from the first week.

The second is that it corrupts its own number. If usage is the metric, usage is what people optimise. You get AI-in-the-loop for tasks that didn't need it, prompts written for the log rather than the outcome, and a healthy chart. You have manufactured exactly the performative usage that was already fooling you, except now it's on purpose.

What I think the alternative looks like#

Instead of policing usage, give each person a way to see whether their own work is actually moving. In their own language, on their own material, with AI as one of the inputs rather than the thing being counted.

The unit of measurement changes. You stop asking "how much AI is being used" and start asking "did the work get better, and where." AI shows up in that picture as context, not as the score. Adoption stops being a proxy for performance, which is good, because it was always a bad proxy.

That's a design position, not a feature list. But it's the position that follows from the three findings above, and I haven't found a way around it.

Which is where I have to say what I can't do, because it's the same sentence as the argument.

The most requested slide in this market right now is the one that proves the AI spend paid off. I'm pre-launch and I can't produce it for you. More to the point, every version of it I've been shown so far falls apart on the second question, and the second question is always some form of "compared to what." If someone hands you that slide, ask them. I'd genuinely like to be wrong about this and I'd want to see the one that holds.

So I'm working the smaller claim first: one person, their own work, enough signal to tell whether it's actually moving. That's the layer running today. The team-level answer has to sit on top of it, not replace it, because an aggregate built on numbers nobody trusts is just a faster route to a confident wrong answer.

One question back to you#

If you've mandated AI: what did you tell your board it bought you?

I'm asking because I suspect writing that sentence honestly is harder than the mandate was, and I want to know how other founders actually phrased it.