日本語 Get in touch
All articles

Productivity

Developers were 19% slower with AI and felt 20% faster

In a randomised controlled trial with experienced developers, working with AI tools took 19% longer. The same participants reported that AI had made them 20% faster.

Published 2 min read

Knowing this study makes internal conversations about AI effectiveness more accurate.

METR published a randomised controlled trial on 10 July 2025. Experienced open source developers worked on real issues in repositories they already maintained, with AI tool access randomly assigned.

Tasks in the AI condition took 19% longer to complete.

The same participants reported that AI had made them 20% faster.

How to handle this study honestly

Using this as evidence that AI does not work would be wrong. The conditions were specific.

Participants were expert developers working in codebases they knew intimately. That is the setting where AI's relative advantage is smallest. When the whole system is already in your head, the cost of receiving and evaluating external suggestions can exceed the benefit.

An unfamiliar language, a codebase you have never seen, routine implementation work. Those conditions could produce a different result.

METR themselves published a partial revision of their original interpretation on 24 February 2026. Citing both halves is the honest treatment, and that honesty is part of what makes the study worth citing.

The important number is not 19%

What matters operationally is not the 19%. It is that people who got slower believed they had got faster.

The gap is 39 points.

Which means the effect of AI adoption cannot be assessed by asking how it feels. In this study, the practitioners' own impression pointed in the opposite direction to the measurement.

Many companies evaluate AI rollouts with a staff survey asking whether work has become more efficient. This study suggests that question may not recover the truth.

Why it feels faster

This part is inference, but the quality of the waiting appears to change.

Time spent thinking registers as effort. Time spent waiting for a generation does not register the same way. Even when total elapsed time is identical, the work feels less tiring.

Less tiring reads as faster. That is not an illusion. The experience genuinely improved. It is simply a different variable from time.

A report that AI made the work easier is usually accurate. A report that it made the work faster should not be believed until it is measured.

Decide the measurement before you start

Three things to settle in advance.

  1. Get time from something other than self report. Ticket opened to closed. Pull request opened to approved. Use a metric the system records on its own.
  2. Watch quality at the same time. Speed alone hides an increase in review burden. Track rework rate, review comments, and post release fixes.
  3. Ask for the subjective view separately. Perception is worth collecting. Keep it in a different column from the time data.

Japanese generative AI usage has reached 86.4%. From here the question stops being whether to adopt and becomes what adoption produced. At that point, a company with no measurement plan will decide based on how it felt.

This study is a record of how being wrong by 39 points looks.

Sources

  1. METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (10 July 2025, with a partial revision published 24 February 2026)
  2. MIC Japan, White Paper on Information and Communications 2026 (published 2026-07-24)

Recognise the problem

We design, build and hand over the layer that turns a stalled AI pilot into something people actually use. Tell us where you are stuck.

Get in touch