WRITING / POST
The productivity is there. Organisations can't collect it yet.
The productivity problem
The gains are real where the work is done, however organisations can't collect them yet. AI made the work faster, but the larger the organisation, the larger the inertia of existing processes, and AI hasn't overcome that or made the organisation faster.
On 5 October Torsten Slok, chief economist at Apollo, published a chart under the title "No Signs of AI in the Productivity Data". Output per hour is running at about 2.5%, which is healthy. Utilisation-adjusted total factor productivity, the measure that ought to move if a new technology is making everything work better, is sitting slightly below zero with no sign of acceleration. His reading is that strong output per hour with flat TFP looks like capital deepening, firms buying more kit, rather than a technology shock, and that the AI payoff remains "a forecast rather than an observation". He also notes that electricity and IT each took a decade or more to show up.
In my view, he is right about the data, but is too quick to read flat TFP as absence.
The disconnect
I had three, and now four, agent 'teams', AI agents that specialise in my own production tasks, and at the same time double as case studies and experimentation in real world agentic deployment. This article itself was written with the considerable assistance of my editorial agent, Clare. Whether this and other articles are of utility to anyone remains for the reader to decide, but the grammar is better and all the typos have been corrected at least.
It does not surprise me that one professor has produced 200 academic papers in nine months. Are the papers good, or just an agentic slop mill? Well, if his agentic editor assistant is anything like mine, it won't let him get away with unsubstantiated claims or anything that won't stand up under peer scrutiny.
Where the gain stops
The best evidence I have seen on where the gain goes is a working paper by Fiona Chen and James Stratton at Harvard, Artificial Intelligence in the Firm: Bottlenecks in Software Production. They used data from Jellyfish, an engineering analytics platform, covering 718 firms and some 300 million work events across GitHub, Jira and calendars, and compared firms according to when they adopted AI coding tools.
When firms adopted coding agents, lines of code went up 30%, commits 20% and pull requests 23%. The number of Jira issues and epics resolved, which is much closer to what the business actually gets, did not move by any amount they could detect, and they can rule out an increase larger than about 12%. Nor did the issues get bigger to soak up the extra code.
The gain stopped at code review. Time from submitting a pull request to merging it rose 49%, from about seven days to ten and a half. The share of pull requests sent back for changes nearly doubled, from 13% to about 25%, and reviewers left 35% more comments on pull requests that were no bigger than before. Firms responded by drawing about 14% more of their people into review, without taking anyone off coding. Most have tried AI review tools, but AI wrote only 23% of review comments. Humans are still doing the checking, and there is more of it to do.
That is the part of the gain that gets paid back. The authors separate two effects of AI: more code, and a higher rate of problems in that code. A team can add reviewers to cope with more volume. It cannot add its way out of code that needs more checking per line. One of their interviewees put it plainly: "The AI tools lack the comprehension of the overview of the entire feature, and what it is supposed to do."
Adoption is also not the same as use. Over 95% of the firms in the sample had adopted coding agents by January 2026, but twelve months after adoption only about 20% of engineers at the median firm were using them, and in the bottom fifth of firms almost none were. We have all heard the anecdotal cases of Microsoft Copilot being switched on, then switched off again six months later because nobody could see a benefit.
Napoleon's Maxim XX has it that "the line of operation should not be abandoned; but it is one of the most skilful manoeuvres in war, to know how to change it". I know first hand that it is really, really hard to change a company's line of operation, and when you are bold enough to try, the dildo of unintended consequences never arrives lubed.
The paper has limits. It is a working paper, not yet peer reviewed. It covers software only, the firms are Jellyfish clients and so probably keener on measurement than most, and the window is about a year. The effects are for agents; the earlier autocomplete-style assistants barely moved anything.
Chen and Stratton's prescription is complementary investment in review capacity: more reviewers, people moved across from coding, better review tools. Where the problem is the quality of the code rather than its volume, they concede the fix has to come from the technology itself. I come from the era of rough consensus and running code, and I don't think either investment or technology in the above context gets there. The existing pipeline can't stand, nor can better technology, as it is today. To properly account for the interdependencies, and not create more issues than it resolves, I think it will take a model with roughly a thousand times today's context, and attention that genuinely holds across all of it. That's my guess, not a measurement.
Why the statistics are flat
None of this is new to economists. Humlum and Vestergaard, using Danish administrative data, found widespread AI adoption and worker-reported benefits alongside precise null effects on earnings and recorded hours. Their explanation is that employers "absorb AI through task reorganization", and that "technological change reshapes work well before it surfaces in earnings or hours." The Jellyfish data shows that absorption happening at the point of production.
Brynjolfsson, Rock and Syverson's Productivity J-Curve (2021) makes the general case. A technology like this needs large investment in things the national accounts don't count: new processes, new skills, new ways of organising work. While that investment is being made, measured productivity is understated, and it is overstated later when the payoff arrives. Flat TFP now is what the model predicts.
Paul David made the historical case in 1990 in "The Dynamo and the Computer", and it is the reason behind Slok's own electricity caveat. Factories first electrified by swapping the steam engine for a large electric motor and keeping the same line shafts and the same floor layout. The productivity came decades later, when factories were rebuilt around a motor on each machine. Patching the old layout with the new power source was not the answer then either.
There is also a mechanical point. Firms get AI by buying compute and model services, and growth accounting books those as capital. Some of AI's effect therefore turns up as capital deepening by construction, which is exactly where Slok sees it.
The organisation of one
My own use stands testament to the productive multiplier of AI. Two years ago my output was infrequent maintenance of my gemstone eCommerce site, a pretty poor effort at a personal web site, and writing a single article of interest, or a substantive Quora answer about once a month. Over the last year, through integration of AI, I have fully functioning system administration, 2-3 substantive blog posts a week, three technical papers, a planning prototype for mine site development, released 50 music tracks across five concept albums, and am in the process of handing off gemstone promotion to a new AI agent. All documented here, documentation being another thing AI has enabled. And I haven't given up my day job of chief cook and bottle washer around the house either.
That increase in output came at a cost, true. It has not been without considerable frustration. I could collect the dividend because I built the workflows for my own specific methods. It was a steep learning curve, and that is where the bulk of the frustration came from: an agent that made the monitor green by removing the check, a system that worked while broken, and a harness that loaded skills it had no business following, are a few examples, if you care to know.
Nevertheless, the gain is undeniable, the price is now paid, and there is no going back.
What this doesn't prove
In July 2025 METR published a randomised trial with 16 experienced open-source developers working on 246 tasks. With early-2025 tools they were 19% slower, while believing they were 20% faster. Felt gains can exceed real ones. My own evidence is a count of things shipped, not a sense of speed, but whether they're worth anything is a separate question. The tools have moved a long way since (METR itself now calls those results out of date), and the trial measured how fast known tasks got done rather than whether new work became possible, but it is a fair warning.
Not all of the gap is inertia, either. If there is only so much work to do, freed capacity has nowhere to go, and that is a limit on demand rather than a failure to organise. And some of the slowness is deliberate. Firms keep humans in review because shipping bad code is expensive, and that is a judgement about liability, not a lack of imagination.
Where this goes
My view is that organisational structure optimised for efficiency and productivity in a pre-AI world will not be the structure that works in an AI-embraced world. And I don't have any answer to what will, though I have an emerging shape of what I think it will be like. What I think the shape is, is that each individual will become a hub of their own industry within an organisation geared to direct that output towards its goals.
Takeaway summary
- AI productivity realisation points to needing fundamental change to organisational structure.
- AI enables greater output for the individual.
- Realising AI potential will come from re-engineering around the AI/human symbiosis.