AI Made Writing Code Free and Moved the Bottleneck to Code Review

Sep 29, 2026
•
8
min read
Team 1up
Team 1up
AI Made Writing Code Free and Moved the Bottleneck to Code Review
Share this post

Drafting used to be the slow part. Now an agent opens a pull request in ten minutes, and a chat tool produces a full first pass at a security questionnaire in seconds. That is a huge win, and it moved the interesting problem one step down the line, to the person who reads the output, trusts it, and puts their name on it. Four engineers at 1up raised this on their own during a round of interviews about AI, and none of them were asked about code review. It maps to the AI trust gap that most teams are working through right now. Speed arrived first. Confidence is catching up, and there are good ways to speed that up too.

Key Takeaways

  • AI removed most of the cost of producing a first draft. That success shifted the attention to human review and approval, which is a very solvable next problem.
  • Review time has grown alongside real gains. Telemetry from 22,000 developers shows median time in pull request review up 441.5%, next to a 66% jump in epics completed per developer.
  • Reviewers are doing more valuable work than before. AI-written code reads well on the surface, so review is now about intent and behavior rather than syntax.
  • Output and review capacity grow best together. Pull requests merged with no review at all are up 31.3%, which points to a queue that needs support.
  • Teams getting great results are reshaping the work. Smaller changes, written standards the agent follows, and a clear record of what the system does all lighten the load on reviewers.

Why human review became the bottleneck in AI-assisted engineering

One engineer at 1up started building a side project this year for a single reason. The agents got faster than he could read.

"Generating code has never been easier with AI but now code review becomes the bottleneck. AI can introduce all this new code and features but how does one review it all, especially newer unfamiliar features? I think there will be a shift away from reviewing just the code and we will focus more on the Product Features and Behaviors as well as architecture" said Edward Poon, Co-Founder and CTO at 1up.

A pull request, or PR, is the change a developer submits for a teammate to check before it gets merged. Coming from the person who runs engineering at an AI company, that comment is a sign of progress. The tools are working well enough to reveal the next constraint, which is how good systems improve. Faros AI found that 80% of the teams it studied now pass the 50% weekly active user mark for AI tools, and the acceptance rate for AI-generated code has climbed from 20% to 60%. DORA's 2025 research puts adoption higher still, with 90% of technology professionals using AI at work and more than 80% saying it made them more productive.

So the code arrives, and it arrives in volume. The number of people available to approve it stayed about the same. That gap is the thing worth designing around, and teams are already doing it.

The data on AI code review time shows where all the work went

Two things are true at once. Output went up. So did everything sitting behind it.

Review metrics under high AI adoption, measured by Faros AI across two years of telemetry from 22,000 developers and more than 4,000 teams:

  • Median time in pull request review: up 441.5%
  • Average time spent in code review: up 199.6%
  • Median time to first review: up 156.6%
  • Pull requests merged with no review at all: up 31.3%

That last number is the most useful one. Faros reads it as a capacity issue rather than a shortcut anyone chose. Reviewers are handling as much as they can, and the overflow finds its way through. Capacity problems respond well to better tooling and better scoping, so this one is worth being optimistic about.

The same research shows the upside clearly. Epics completed per developer are up 66%, task throughput per developer is up 33.7%, and pull request merge rate is up 16.2%. Downstream, code churn is up 861%, bugs per developer are up 54%, and incidents per pull request are up 242.7%.

Bugs per developer sat at 9% in the 2025 version of that research and reached 54% a year later, which is a good early warning to act on now. Google's 2025 DORA report found the same tension from another angle, with AI now showing a positive relationship with throughput and a negative one with delivery stability. The speed is real and it is holding. The systems around it are still catching up, and that part is within any team's control.

Teams outside engineering will recognize the shape of this. It is the same reason wrong answers from internal AI tools tend to surface later than anyone would like, and the same habits fix both.

Source: Faros AI

What the workload looks like when 90 percent of code is written by AI

One of the founders put a number on how much the writing itself has changed.

"Ninety percent of our code is written by Claude. There is no more manual writing. It is guiding, reviewing, and high level architecture. Which means there is a lot more code to review," said Ofer Shakarov, Co-Founder and Lead Software Engineer at 1up.

He also described a specific kind of fog that comes with that volume.

"You see a bunch of files sitting next to each other and it is not always clear what is happening," Shakarov said.

Then he named something a lot of engineers feel and rarely say out loud.

"The manual part used to be a break. You would sit down, think about what you were going to do, then write it out. Your brain got a rest during the writing. Now only the creative part is left, back to back," Shakarov said.

That matches the research closely. DORA's qualitative work found that time saved during creation often gets reallocated to auditing and verification. The hours moved into higher-leverage work, and a smaller group of people absorbed them. Faros calls this the senior engineer tax, and it shows up as daily pull request contexts per developer up 67.4% and 26% more in-progress tasks with no activity for a week or longer.

More gets started. Pacing is what turns that into more finished, and pacing is a design choice a team can make.

Reviewing AI-generated code can cost as much as writing it yourself

The most direct version of the observation came from the engineering side of the team.

"Reviewing it sometimes takes as much time as writing it myself," said Nick Weaver, Software Engineer at 1up.

He is genuinely positive about the trajectory and said agentic coding has improved a lot over the last nine months. His note is about the shape of what comes back.

"It gives you overly complicated solutions. It would be more elegant if you just did it yourself. It tries to solve every problem like you are a huge company, so there is a lot of overkill," Weaver said.

He summed up the day-to-day experience this way.

"Some days it is so awesome. Other days I hate it and it is unreliable," Weaver said.

Overkill shows up most at review time. A change that should touch two files touches nine. Faros measured pull request size climbing as adoption grew, which adds reading time to every approval. Tighter tickets and narrower scopes give a surprising amount of that time straight back.

Why AI-generated code is harder to review than human-written code

There is a good reason this feels different from a normal backlog.

AI-written code usually looks clean. Good names. Consistent style. Sensible structure. The structural and logical gaps, when they exist, sit beneath that surface. Catching them means reading carefully, reasoning about intent, and reconstructing the problem the code was meant to solve. That is deeper work than scanning a diff, and it is worth building the review process around.

"With a big code base it is hard to keep track of everything, so it is hard to catch slop. The slop existed before AI. There is just more of it now," said Shakarov.

He also gave the best analogy of the four interviews.

"You can trust an AI to design a closet in your bedroom and then build it yourself and it’s good enough. I would not trust AI to scale that to a whole apartment complex. Not yet. Eventually it might get there," Shakarov said.

The gap between a shelf and a building is taste, and taste is where humans still add the most value. Anyone working through AI slop in writing knows the feeling. Output that looks right deserves a closer read, and building that read into the process pays for itself.

How engineering teams are working around the AI review bottleneck

None of the four engineers wants to go back. They reshaped the work instead, and the changes are practical enough to copy.

Anchor review to behavior instead of code. One code review tool built in-house maps the code’s call stack, then hands the reviewer the new or changed behaviors.

"I am inundated with 30+ Merge/Pull Requests that are 1,000+ lines long, it gets very tiring to review it all. We are making a trade off where we assume the AI wrote syntactically correct code but the issues lie more in the architecture and product behaviors. I care about what it is doing. When I add a connector, I need to know if it syncs that space with our knowledge base. In the past I would read the code to find out, or ask Claude and hope it grepped the right text but even that can hallucinate and it also takes forever," said Poon.

The payoff is a small input instead of an enormous one.

"Rather than reading disparate files and mentally piecing it together, I look at the changed product behaviors and architecture." Poon said.

He calls the step-by-step paths through the code "walks," and he feeds those back to the model so it stops hunting.

"Then it is not searching around for twenty or thirty minutes trying to figure out the recipe. It already has the recipe," Poon said.

His forecast follows from that.

"The rate in which code is changing due to AI forces folks to update their mental model of the codebase constantly. It becomes infeasible to hold the specific lines of code in one’s head. With that said, we should be anchoring on product behaviors and architecture and not focus on the code as much because AI generally writes syntactically correct code. That is where this is going," Poon said.

Write the standards down where the agent will read them. The team keeps its coding rules in one file that every agent has to follow.

"When someone joins, we define the standards. The patterns for our code base that have to be followed. Then we turn that into a guide any AI agent has to follow. We keep it in a markdown file at the root of the repository. We needed it, because otherwise the model just builds whatever it thinks is right," said Antonio Carlos (Tonni) de Andrade d'Avila Pacca, Senior Software Engineer at 1up.

Rules at the front save time at the back. It is the same thinking behind AI governance for enterprise knowledge, applied to a repo instead of a knowledge base.

Keep the skills that make review work. One habit on the team is worth borrowing.

"Sometimes we need to stop using the AI and go back to writing code ourselves, otherwise we forget how. So every so often I stop and figure something out on my own. I was doing this before AI. I should still be able to do it," Pacca said.

Review quality depends on reviewers who can tell good from plausible, and staying sharp keeps that judgment strong. Roughly 30% of developers in DORA's 2025 survey place little or no trust in AI-generated code, and that healthy skepticism pays off when the person holding it can back it up.

Shrink the unit of review. Smaller pull requests, clearer tickets, one behavior at a time. None of this is new advice. It simply matters more now, and it works.

The same bottleneck is already showing up for teams who do not write code

Swap pull requests for questionnaire responses and the story is nearly identical.

An AI answer engine can draft two hundred answers to a security questionnaire before lunch. Then it lands on a security lead, a legal reviewer, and two subject matter experts who have day jobs and no idea they are on the hook. The draft took four minutes. The approval takes three weeks. Volume went up. Approval capacity did not. That is why working with SMEs on RFPs is mostly a scheduling and trust exercise, and why small changes there produce outsized wins.

The failure mode is the same one Faros measured in code. When a reviewer gets two hundred items that all look equally finished, nothing tells them where to spend their attention. So they skim, or they stall, and the answers that actually needed a second opinion slip through next to the ninety that were fine.

The fixes rhyme with the engineering ones:

  • Show where every answer came from. A source document and a date turn a judgment call into a quick check.
  • Rank the queue. Flag the ten answers that need real scrutiny instead of asking for two hundred equal approvals.
  • Keep a record of what was already approved. Nobody should re-argue the same SOC 2 answer every quarter.
  • Route items to the person who owns them, with a due date and a reminder, so review stops depending on who notices the Slack message first.

Human judgment matters more here than it did two years ago, and it is worth spending carefully. Halving the number of answers a reviewer has to touch beats another speed gain in drafting every time.

Where to start if review is the slowest part of your week

Four engineers, four different jobs, one shared observation. That is a strong signal, and it points to four changes a team can make this month.

Measure the wait, not the output. Merge counts will look great right now. Time to first review and time in review are the numbers that show whether the work is actually moving. Pull both for the last two quarters and compare them against your AI adoption curve. If the wait is growing faster than the output, adding more generation capacity will make the queue worse.

Give reviewers less to read. A pull request that touches nine files to solve a two-file problem costs somebody an hour. Cap the size of a change, one behavior per request, and send the reviewer a plain-language summary of what the code does before they open the diff. Poon's "walks" approach is one version of this. A well-written PR description is a cheaper version that works today.

Put the rules where the model reads them. A standards file at the root of the repo turns tribal knowledge into something the agent follows on every run. Every convention written down there is an argument nobody has to relitigate in review.

Protect the reviewers. They are the reason all that new speed turns into shipped work. Rotate the load so the same two senior people do not absorb every approval. Give them time that is actually blocked off for review instead of squeezing it between their own tickets. And back Tonni's habit of building things by hand now and then, because judgment is the one skill the whole system depends on.

Speed is the part AI already solved. Review capacity is the part a team still controls, and the teams that build for it get to keep everything the tools just gave them.

FAQs

It is what happens when AI tools generate code faster than humans can check it. Drafting speeds up, review capacity stays flat, and work builds up in front of approvers. Faros measured median time in pull request review rising 441.5% under high AI adoption.

‍

It usually looks correct. Clean naming, consistent style, familiar structure. Gaps tend to sit in the logic rather than the surface, so reviewers read closely and reconstruct intent instead of scanning for obvious errors.

At the individual level, yes. Epics per developer are up 66% and task throughput per developer is up 33.7%. Delivery stability is the area to watch, with incidents per pull request up 242.7% and bugs per developer up 54%.

‍

Keep changes small, write coding standards into a file the agent follows, give reviewers behavior summaries instead of raw diffs, and track review time as a core metric. Keeping manual skills sharp helps too, since reviewers need strong judgment about what they are approving.

‍

Related Reads

The Payoff of Connecting Questionnaire work to your Sales Pipeline

17 Sep 2026
•
5
min read
Read blog

Shadow AI: What Your Reps Are Pasting Into ChatGPT Right Now

10 Sep 2026
•
7
min read
Read blog

How to Write an RFP Response That Survives AI Scoring

07 Sep 2026
•
8
min read
Read blog

SIG vs CAIQ vs VSA: A Simple Guide to the Big Three

19 Aug 2026
•
9
min read
Read blog

Why Most Internal AI Assistants Get Built and Then Abandoned

04 Aug 2026
•
8
min read
Read blog
Table of contents

1up your sales team

See a demo of how 1up automates answers in seconds.
Book a Demo