Unity Horizon
← Journal

Applied AI

Does AI make developers faster? What the three most-cited studies measured

A Copilot trial found developers 55.8% faster, a BCG trial found 25.1%, and a study of experienced developers found them 19% slower. Checked 11 August 2026.

Unity Horizon5 min read

It depends whose productivity got measured, and on what kind of task. A trial of 95 freelance developers found AI made them 55.8% faster. A trial of 758 BCG consultants found gains up to 25.1%, and only on tasks the model could already do. A trial of 16 experienced developers found them 19% slower, and the follow-up experiment broke before it produced a usable number. The disagreement is not measurement error. It is who was in the room and what they were building.

Checked against the papers themselves, not a secondhand summary of them, on 11 August 2026.

What did the "55% faster" Copilot study actually test?

Ninety five freelance developers, recruited on Upwork, split into a GitHub Copilot group and a no-tools group and given one task: build an HTTP server in JavaScript, from a blank file, according to the randomized controlled trial published on arXiv. Seventy finished. The Copilot group averaged 71 minutes, the control group 161 minutes, which is the 55.8% figure everyone quotes. It is real. It is also a single greenfield task with no existing code to read and no legacy convention to follow, on a well-known problem the model had seen thousands of times in training. The 95% confidence interval on the result is 21% to 89% faster, wide enough that the true effect could be less than half what gets repeated. That does not make the study wrong. It makes it a measurement of one thing: how fast a stranger writes a small program from nothing.

What did the BCG consultant study find, and why does it matter more?

In the study Dell'Acqua and colleagues ran with Harvard Business School and BCG, 758 BCG consultants, about 7% of the firm's individual contributor staff, were split three ways: no AI, GPT-4, and GPT-4 with a short prompting guide. On tasks inside GPT-4's reach, the AI group completed 12.2% more of them, worked 25.1% faster, and produced work rated 38 to 43% higher by blind human graders. On a task built to sit past what the model could reliably do, the same AI group did 19 percentage points worse than the group with no AI at all. The paper, published in Organization Science, calls this the jagged frontier: capability that helps enormously on one task and hurts on the next, with nothing to tell the user which side of the line they are on. The gain was not evenly spread either. Below-median performers gained the most. The seniority gap is not noise. It is the finding.

What happened to the study that found AI makes developers slower?

Sixteen experienced open-source developers, working in their own repositories, on 246 real issues, with an average of five years' prior experience on those specific projects, were the subjects of METR's 2025 trial. With AI tools allowed, they took 19% longer to finish, with a confidence interval of 2% to 39% slower. It became the standard counter-example to every vendor productivity claim, and it held up under scrutiny the Copilot study never faced.

Then METR tried to repeat it. The August 2025 follow-up recruited 57 developers, ten of them returning, and could not produce a reliable number. Between 30% and 50% of participants admitted avoiding tasks assigned to the no-AI condition, or skipping the ones they thought AI would help with most, so the sample stopped being random. Cutting pay from $150 an hour to $50 likely made the selection effect worse. The study meant to measure whether AI speeds people up ended up getting swallowed by the fact that its subjects would not work without it. For the returning subset, the point estimate was still a slowdown, about 18%, but the confidence interval now ran from 38% slower to 9% faster, wide enough to include no effect at all. METR's own update, published 24 February 2026, says the true effect could be larger than either number shows, and that the experiment is being redesigned. A control group that will not work without AI is not a control group.

So does AI make developers faster, or not?

Study Who was tested What they built or fixed Headline result
GitHub Copilot RCT 95 freelance developers via Upwork, 70 finished One HTTP server, from a blank file 55.8% faster, 95% CI 21% to 89%
BCG consultants (Dell'Acqua et al) 758 consultants, about 7% of BCG's individual contributor staff 18 realistic tasks, some inside GPT-4's reach, some past it 25.1% faster and 12.2% more tasks done inside the frontier, 19 points worse outside it
METR developer RCT 16 experienced open-source developers; 57 in a follow-up that broke 246 real issues in their own repositories 19% slower originally; the follow-up was too contaminated to trust

Inside a frontier that tracks how novel and self-contained the task is, yes, reliably, and the largest gains go to the people who know the least. Outside it, on work embedded in a codebase with years of history and conventions nobody wrote down, no study here shows a reliable gain, and the best-designed one shows the opposite sign with a shaky number attached. Google's 2025 DORA survey of close to 5,000 engineering professionals found AI adoption at 90% and delivery stability getting worse, not better, at teams that added AI without fixing how they ship. AI does not fix a team. It amplifies whichever one it is handed to.

A vendor licensing every seat in the company is selling the 55.8% number to people who look more like METR's experienced developers than Upwork's freelancers. A team scoping AI at one bounded, unfamiliar task at a time is buying the study that applies to them.

None of these three studies ran on a codebase anywhere near the size or the age of what most companies carry, so how the frontier looks after ten years of accumulated code is still assumed rather than measured.

This is also the argument for scoping before licensing: decide, task by task, where the frontier sits for your own codebase, rather than buying a seat for everyone and hoping. It is why we scope builds senior led: the person deciding whether a task sits inside the frontier is the same person writing the code. Map where that frontier sits for your team before the next seat renewal, not after.

Which task on your roadmap this month is genuinely greenfield, and which one only looks that way because nobody on the team has opened that file in years?

Common questions

Does AI make developers faster or slower?
It depends on the task and the developer. The best-designed trials show large, reliable gains for freelancers and junior consultants doing bounded, unfamiliar work, and no reliable gain, possibly a loss, for experienced developers working inside a codebase they already know.
What did the GitHub Copilot '55% faster' study actually measure?
95 freelance developers building one HTTP server from a blank file, with Copilot cutting the average time from 161 minutes to 71. The 95% confidence interval was 21% to 89% faster, and the task had no existing codebase to work inside.
Why is the METR study on AI slowing developers down being redesigned?
METR's 2025 trial of 16 experienced developers found AI use added 19% to task time. A 2025 follow-up could not reproduce a reliable number, because developers refused to be assigned to the no-AI condition, so METR's 2026-02-24 update says the experiment needs a new design.
Does AI help junior developers more than senior developers?
Yes, according to the BCG consultant study, where the largest productivity gains from GPT-4 went to below-median performers. The same study found consultants using AI on tasks past the model's reliable reach did 19 percentage points worse than consultants with no AI at all.

[ NEXT ]

Want this built for you?

Thirty minutes, no prep, no pitch. Tell us what slows your business down and we'll show you what an agent can do about it.

Book a call