Ask anyone who has spent the last two years wiring an AI assistant into their workday whether it has made them faster, and the answer is almost always yes, delivered with total confidence. The best evidence available does not agree with itself, and worse, it suggests that confidence is exactly the thing not to trust.
Start with the result that made AI look like the productivity story of the decade. Researchers at Harvard Business School, working with the Boston Consulting Group, gave 758 BCG consultants a set of 18 realistic business tasks and split them into groups with and without access to GPT-4. For tasks that sat within what the model was actually good at, the AI-assisted consultants were 12.5 per cent more likely to complete the task successfully, finished it 25.1 per cent faster, and produced work independent evaluators rated as roughly 40 per cent higher in quality. Those are not modest numbers. They are the kind of result that gets cited in every AI strategy deck written since.
The frontier is jagged, and stepping off it is expensive
The same study, led by Fabrizio Dell’Acqua and colleagues, also found the part that rarely makes the slide deck. When consultants used GPT-4 on tasks that sat just outside the model’s actual capability, a distinction the researchers call the “jagged frontier” because AI is unpredictably brilliant at some tasks and unpredictably weak at adjacent-looking ones, performance did not just fail to improve. It got worse than the group that used no AI at all. The consultants tended to accept plausible, confident-sounding AI output without checking it closely enough, a pattern the paper’s authors nicknamed falling asleep at the wheel. The tool did not fail loudly. It failed quietly, and the human in the loop did not notice in time.
That would already complicate the simple story that AI makes knowledge work faster. A second study, published by the AI evaluation organisation METR in 2025, complicates it further, this time for a task where confident predictions of an AI productivity boom have been loudest: software engineering.
Slower, and certain they were faster
METR ran a randomised controlled trial with 16 experienced developers, each a maintainer on a large open-source project, working through 246 real coding issues, bug fixes, features and refactors on codebases they already knew well. Each issue was randomly assigned to be done with AI tools allowed or not. The developers used Cursor with a frontier Claude model, and were paid $150 an hour to work as they normally would, screen recorded throughout.
Before the study, the developers expected AI assistance to speed them up by 24 per cent. What the trial actually measured was the opposite: tasks done with AI took 19 per cent longer to complete. The gap between expectation and result is large enough on its own. What happened next is the more useful finding. After the study, having lived through the slower outcome themselves, the developers still estimated that AI had sped them up, by about 20 per cent. Direct, recent, personal experience of a real slowdown did not update their belief that it had been a speed-up.
METR’s own analysis does not settle on one tidy explanation. Among the factors it points to are that most of the developers had only around 50 hours of experience with the specific tool by the time the study ran, and that the review standards on mature open-source projects are unusually exacting, the kind of environment where a plausible-looking suggestion needs more scrutiny, not less. Neither factor makes the result an indictment of AI coding tools in general. Both suggest the gap between feeling faster and being faster is not a fluke of one bad study design.
What the two findings have in common
Put next to each other, the BCG and METR results are not actually a contradiction. They are the same warning from two different angles. AI assistance can produce a real, measurable, sizeable productivity gain, the consulting study’s numbers are large by the standards of any workplace intervention, but only inside the boundary of what the tool is genuinely capable of on that specific task, and only when the person using it is scrutinising the output rather than accepting it. Once a task drifts outside that boundary, or once the human reviewing the work relaxes because the last ten suggestions were good, the same tool can make things measurably worse while still feeling helpful.
Broader industry data, gathered together in one recent review of surveys and datasets, points the same way. A DX survey of around 121,000 developers at more than 450 companies found 92.6 per cent using AI coding tools at least monthly, a figure close to the 90 per cent adoption reported separately in the 2025 DORA report on software delivery. Adoption, in other words, is close to universal. Yet a review of six independent studies on organisational productivity found the gains converging on roughly 10 per cent at the level of a whole team or company, nowhere near the 25 per cent the BCG consultants achieved on tasks squarely inside the frontier. One dataset from Faros AI, tracking more than 10,000 developers, found teams with the heaviest AI use merged 98 per cent more pull requests, while review time rose 91 per cent, reported bugs rose 9 per cent, and the broader delivery metrics organisations actually care about stayed flat. Near-universal adoption and a large, well-documented capability boost have not, so far, translated into a large measured gain once the work is checked all the way through.
That last part is the awkward one for anyone hoping to manage this with a simple rule. Feeling faster turned out to be a poor guide to being faster in the one study that checked. The BCG consultants who wandered off the frontier presumably also felt they were being helped, right up until an evaluator scored their output. None of this argues for abandoning these tools, the productivity gains inside the frontier are real and worth having. It argues against trusting your own sense of how well a session went as the measure of whether it actually worked, which is precisely the measure most people are relying on.
I wrote a while back about how being busy and being productive are routinely confused, and this feels like the same mistake wearing a new tool. The confidence that something worked is cheap to produce and easy to feel. Whether it actually did is a separate question, and on the evidence so far, it is one that generally needs an outside check to answer honestly.