The AI Agent Did Exactly What You Asked, and That's the Problem
When an autonomous agent turns in weak work, the reflex is to fault the model. The problem often sits upstream in the human’s acceptance criteria, making the ability to define outcomes without prescribing steps increasingly scarce.

In my experience building autonomous development agents, I found that whenever I had to steer them mid-task, the issue was not the agent's capability. Every single time it was something lacking in my acceptance criteria.
If this caught your attention, that’s not accidental.
The best editorial systems don’t happen by accident. Outlever builds them.

An autonomous coding agent starts a task, and a few steps in it drifts. It solves the wrong version of the problem, or takes a clumsy route to the right one, and the engineer watching reaches for the obvious explanation: that the model just isn't good enough yet. Liron Tzabari, Director of Engineering at DoubleVerify, an ad-verification and media-quality technology company, thinks that explanation is usually wrong. He wrote on LinkedIn that we're too fast to fault the tool. "Can we really blame the AI Agent for poor results?" he asked. "Often, we judge performance by a snapshot rather than the curve."
What he found building the agents is less flattering to the humans running them. "In my experience building autonomous development agents, I found that whenever I had to steer them mid-task, the issue was not the agent's capability," he wrote. "Every single time it was something lacking in my acceptance criteria." The moments he had to grab the wheel weren't the model coming up short. They were the moments his own instructions ran out.
The easy read of an agent going sideways is that it needs a stronger model, a bigger context window, or one more tool bolted on. Tzabari's account moves the fault the other direction, back to the brief the human wrote before the agent ever started. If the target was fuzzy, the agent optimized a fuzzy target, and it did so fast enough that the gap only surfaced once the work came back wrong.
Define the outcome, not the route
He draws the line between two ways of asking. "When I prescribed solutions, agents did what they were told, but the results were suboptimal," he wrote. "When I defined outcomes, the agents found better paths." Hand an agent a step-by-step recipe and it runs the recipe, mistakes included. Hand it a clear picture of what a good result looks like and the constraints that result has to satisfy, and it has room to reach an answer you wouldn't have scripted. "You don't tell them what to do. You tell them what you want to achieve and how the result should look."
The spec strikes back
None of that's new to anyone who has shipped software, and that alone is notable. The failure he's describing, work that goes wrong because nobody pinned down what "done and good" actually meant, is one of the oldest in the field. The Project Management Institute's research attributed 47% of unsuccessful projects missing their goals to poor requirements, the largest single cause it measured. A benchmark study from IAG Consulting found that companies with weak requirements practices hit three times as many project failures as successes, and that 68% of them were likelier to land a marginal or failed project than a good one. Long before agents, the spec was already where projects went to die. The agent didn't invent that failure mode. It just runs a flawed spec in minutes instead of quarters, so the cost of a vague brief arrives faster and in plainer view.
That's the shift sitting under the tooling. Agents haven't retired engineering judgment, but they have relocated where it gets spent. The scarce skill stops being who can write the function and becomes who can specify what the function has to be, sharply enough that an eager, literal worker can't satisfy the letter of the instruction while missing its point. Writing crisp acceptance criteria turns out to be harder than writing the code, because it forces you to decide what you actually want before anything gets built. Tzabari reaches for the parallel himself. "Interestingly, is it any different with people than with agents?" he wrote. A manager who tells a junior engineer which keys to press gets compliance and mediocre work. One who defines the outcome and the bar it has to clear gets judgment, and often a result better than the one they'd have dictated.
The engineer who knows what 'good' looks like
For engineering leaders, an agent going off course is useful evidence. It shows exactly where the brief stopped being clear and judgment had to fill the gap. That makes acceptance criteria more than a process skill. They’re becoming a hiring signal. The strongest engineers will be the ones who can turn an ambiguous goal into a definition of success that another person or an agent can act on without constant rescue.
Every mid-task correction reveals something about the worker, but it also reveals who knew how to frame the work.
SiiRA connects US companies with top international talent — end to end, effortless.

The best candidate is not in your city.
SiiRA connects US companies with top international talent — end to end, effortless.

See talent differently.
Get the latest ideas on hiring, leadership, and the future of work.






