Blog
Updates, commentary, and links. No method section, no failure boundary, just the straight read on whatever is on our mind. Want that rigor run on your own question? Here is how commissioned research works.

Developers Using AI Were 19% Slower. They Reported Being 20% Faster.
Adoption is a measurement claim, and fewer than one in five companies own the instrument that would let them make it. The one controlled trial that put a stopwatch next to the feeling found the two disagreed by 39 points, which is a problem for every self-reported AI productivity number in circulation.

A Pipeline With No Agents Beat Every Agent Framework. It Cost 70 Cents a Task.
Teams are buying autonomy they do not need and skipping the structure that demonstrably works. Anthropic's own numbers put multi-agent systems at 15x the tokens of a chat, with token spend alone explaining 80% of the performance gain. Meanwhile the one controlled test of whether scaffolding becomes unnecessary as models improve rejected its own hypothesis.

88% of Companies Use AI. About 5% Get Anything Out of It.
Adoption is nearly universal and value capture is rare, and the gap is not the model. Across McKinsey, BCG, Stanford, and MIT, the one thing that separates the companies getting returns is that they redesigned the work, not that they bought a better model. Here is what the controlled evidence actually supports, and where it runs out.

Your AI Automation Will Outlive the Model It's Built On
GPT-4's accuracy on a simple task fell from 84% to 51% in three months, same vendor, same model name, nobody touched the workflow. The fix is building the workflow so the model underneath it can change without breaking anything.