Switch Edition
Home

>>

Technology

>>

Artificial intelligence

>>

Where Engineering Work Goes Wh...

ARTIFICIAL INTELLIGENCE

Where Engineering Work Goes When AI Writes the Code

Where Engineering Work Goes When AI Writes the Code
The Silicon Review
10 July, 2026
Author: Artem Zamarajev, Senior Software Development Engineer At Amazon

AI can make code much faster to produce, but that does not mean the whole software development process speeds up at the same rate. Much of the work moves elsewhere – into review, rework and the judgment required to decide which changes are worth taking forward. This article looks at what that shift means for engineering productivity and how teams should measure it.

In the time it used to take me to write and merge one pull request, AI agents now produce about ten candidate changes, and I merge roughly six. Code has become cheap to produce, but software delivery hasn't sped up by nearly as much. Much of the effort has shifted to reviewing changes and deciding which ones to keep, and that work is easy to miss when teams count code output.

image

Code Is Cheaper, but Delivery Is Still a Pipeline

A change still has to be reviewed, tested and deployed. Sometimes it also has to be fixed after it ships. AI has shortened the coding step, but the other steps still determine whether the whole pipeline gets faster.

Developers' own sense of speed can be a poor guide, so it helps to check that impression against tracked data. In METR's randomized controlled trial from 2025, sixteen experienced open-source developers completed 246 real tasks in repositories they knew well. They expected AI to make them 24% faster. Afterwards, they felt 20% faster. The measured result went the other way: they were 19% slower (with a confidence interval ranging from 2% to 39% slower). 

A February 2026 follow-up with newer tools estimated that returning developers were 18% faster, but METR said that the estimate was unreliable: 30% to 50% of participants had kept some tasks out of the study because they didn't want to risk doing them without AI, so the tasks where AI was expected to help most were left out.

More pull requests did not mean faster delivery in Faros AI's telemetry, which covered 1,255 teams and more than 10,000 developers. On high-adoption teams, developers merged 98% more pull requests, yet Faros found no significant improvement in company-level delivery metrics. Review shows where the gain went: pull requests grew 154% larger and review time rose 91%. Google's 2025 DORA report is more positive about speed: higher AI adoption was associated with greater delivery throughput, but also with greater software delivery instability, meaning more change failures and rework.

What Changed in My Own Work

At Amazon, I have worked on automation and AI-assisted tooling that catches regressions in code changes before they reach production. None of my own code was AI-generated in June 2024. By June 2025, about half of it was AI-generated, and by January 2026 that share was close to 100%. I still review all of it.

Review is where my work has changed most. I spend less time on style and more on two questions: does the change fit the business need and the existing architecture, and does it use the dependencies and mechanisms the system already has?

The four candidates out of ten that I reject usually go through several iterations without my supervision before I look at them and decide the overall approach is wrong. I spend little active time on those iterations, but I still need enough context to understand why they should be rejected. That context costs time, and no metric I know of counts it, because a rejected candidate never shows up as a merged change. It resembles the pattern seen in the Faros data: more output, with additional work appearing further down the delivery process.

When an agent fails, the cause is usually missing context, not bad code. With incomplete context, an agent can head in the wrong direction or rebuild something the codebase already has, while writing code that looks reasonable. The same happens when it can't reach a system that holds information the change depends on. Giving the agent enough context and checking its plan early prevents much of this, but not all of it.

Sometimes an agent splits the work across subagents: some write code, others review it, and the work goes through a few iterations. At the end, some have built something that makes sense and others have built something useless. Passing another agent's review doesn't mean the work is worth keeping, so deciding whether to keep it is still my job. That sorting and valuation time is part of the cost of each AI-assisted change.

Another effect is harder to measure. In my experience, AI makes it easier to keep a steady pace over long stretches of work, partly because progress shows up faster. Over months, a pace a team can sustain matters more than short bursts. The risk is that at this pace it becomes tempting for engineers to approve code they don't fully understand, which can save time now and cause rework later.

Measuring AI’s Impact

Compare the same team's work before and after adopting AI. A control group working without AI is becoming hard to find, as METR's follow-up showed. Follow similar pieces of work over weeks or months, from implementation through review and rework into production.

Split the results by type of work. AI's effect on implementation or debugging can look very different from its effect on UX decisions, architecture or incident response. A single team-wide number, such as average cycle time per change, can hide those differences. For example, if AI cut implementation time on a team while architecture changes needed more review, the average might barely move. The gains and the extra work both happened – just in different parts of the process.

Elapsed time and effort can also move in opposite directions. With agents writing most of the code, each feature takes less of an engineer's effort, so one engineer can run several in parallel. But attention is split across them, and lower-priority features wait while higher-priority ones get reviewed first, so lower-priority features take longer on the calendar.

What to Track Instead

Pull-request count tells only part of the story. The larger change is where engineering attention goes: generating a candidate implementation now takes much less active effort, while deciding whether it fits the product and the system still requires judgment.

Review time and rework per change should be tracked as well, while effort per feature and elapsed time per feature should remain separate because they answer different questions. Teams should also account for the reviewer time spent evaluating candidates that never merge. No single output metric captures that shift. A rejected change may never appear in delivery metrics, but rejecting it still takes engineering judgment.

Comments

Loading comments…
Loading comments…

MOST VIEWED ARTICLES

RECOMMENDED NEWS

Client-Speak Magazine Subscribe Newsletter Video
πŸš€ NOMINATE YOUR COMPANY NOW πŸŽ‰ GET 10% OFF πŸ† LIMITED TIME OFFER Nominate Now β†’