AI + Technology

AI Can Write the Pull Request in Minutes. The Bottleneck Now Is Deciding Whether to Ship It.

July 21, 2026

Finished-looking code is flooding review queues, raising the value of experienced engineers who can catch regressions, build meaningful evals, and own the outcome.

AI Can Write the Pull Request in Minutes. The Bottleneck Now Is Deciding Whether to Ship It.
Credit: Talent Observer

The interesting work has moved off the model and onto everything around it. The loop. The evals. The judgment about what to trust and ship.

Jason Mattiace

Engineering Head
@
Qualified

A coding agent can now open a dozen finished-looking pull requests in a morning. Each one has green tests, a clean diff, and code that does roughly what the ticket asked. The hard part starts after that, when a senior engineer has to decide whether any of it is actually correct, whether it broke something two modules over that no test happens to cover, and whether it's safe in front of paying customers. The writing was the easy part; the deciding is the whole job now.

The popular version of the AI coding story runs the other way. Generation got cheap, so engineering got easier, so the headcount math gets to shrink. That reading gets the first clause right and everything after it wrong. When a machine takes over the step that used to be expensive, the expense doesn't vanish. It moves to whoever has to stand behind the output.

When code stops being scarce

Jason Mattiace, who runs engineering at Qualified, the conversational-marketing company Salesforce acquired, spent a few days at the AI Engineer World's Fair in San Francisco and came home with that same conclusion in sharper form. "The interesting work has moved off the model and onto everything around it," he wrote on LinkedIn. "The loop. The evals. The judgment about what to trust and ship." And he named where the jam sits. "Code review is the current bottleneck."

The reframe matters because it relocates the constraint. When implementation was the slow, costly step, the scarce person was the one who could turn an idea into working code. Push that cost toward zero and the scarcity moves downstream, to review, integration, and the call about what's safe to run. Google's 2024 DORA report put numbers on the gap. As AI adoption rose, individual work sped up, with measured gains in code-review speed and code quality, and at the same time delivery throughput fell an estimated 1.5% and delivery stability fell an estimated 7.2%. More code, produced faster, did not turn into more software shipped or fewer things breaking. What the report keeps pointing at is everything that happens after the code exists.

Code that looks better than it is

Reviewing that output is getting harder, not easier. GitClear's analysis of 211 million changed lines of code found that the share classified as refactored or reused fell from about a quarter of all changes in 2021 to under 10% by 2024, while copy-pasted code rose enough that developers, for the first time in the dataset, pasted more than they moved. Duplicated, lightly-edited code is exactly the kind that passes a quick read and fails in production. It hands the reviewer more surface to check and fewer of the structural cues that used to make a bad change stand out.

That's the work Mattiace watched get promoted from chore to craft. "Trust became an engineering discipline," he wrote. The energy in the room had moved from "look what the model can do" to "prove it didn't regress in production." Observability, gates, governance, and evals, he wrote, are "core work now, not afterthoughts." None of that comes with a license key. It's built by people who understand the system well enough to know what could go wrong inside it.

Building your own evals

The evals are where this gets concrete for whoever runs the team. The public versions of these tests, Mattiace wrote, have stopped being worth much. "Nobody serious trusts the leaderboards anymore. A model topping a public benchmark tells you almost nothing about how it performs on your actual work. The teams pulling ahead build their own eval sets from real cases." A leaderboard score is a claim about someone else's problem. Your own eval is a claim about yours, and building one takes an engineer who knows which failures actually cost the business something.

For anyone leading an engineering org, that redraws the hiring case instead of shrinking it. If every competitor can generate the same code from the same models, the edge is no longer who writes it. The edge is who can look at generated work and tell what's load-bearing, who can build the eval that catches the regression before a customer does, and who can make the ship-or-hold call and own the outcome. Those are the most experienced people on the team, and this raises their value rather than lowering it. At a company like Qualified, where the product Salesforce bought is an autonomous agent that acts on a customer's behalf, the judgment about what to trust and ship isn't a support function, it's the product.

Generation got cheap. The judgment about what deserves to ship did not, and that judgment still lives in a person who can be held responsible for getting it wrong.

The best candidate is not in your city.

SiiRA connects US companies with top international talent — end to end, effortless.

Hire Top Talent

The best candidate is not in your city.

SiiRA connects US companies with top international talent — end to end, effortless.

See talent differently.

Get the latest ideas on hiring, leadership, and the future of work.

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.