AI + Technology

Writing Code Got Commoditized. Judgment, Design, and Evals Did Not.

July 21, 2026

Any team can watch an agent work. Far fewer can prove the result is any good.

Writing Code Got Commoditized. Judgment, Design, and Evals Did Not.
Credit: Talent Observer

Once a few hundred engineers begin using coding agents every day, the skills, subagents, tools, and rules pile up fast. It's easy to lose track of which ones are pulling their weight. Sprinklr's Director of Product Engineering Rahul Chauhan, who owns the technical direction for the company's content marketing product, watched exactly that happen around his group's Cursor agent. So he had the team build an observability layer on self-hosted LangFuse to measure it, with two interns doing much of the work. "We instrument our services obsessively and then ship them on faith," he wrote on LinkedIn, describing what the project set out to fix.

When someone says writing code is getting commoditized, the reflex conclusion is that engineers are getting cheaper, that a team can be smaller, that the craft is on its way out. But it shouldn't be taken as a budget cue. "Writing code is getting commoditised, engineering judgement, design sense and high agency are not," Chauhan wrote. The commoditization is real, but it applies to writing code, not to the full work of engineering.

You can't benchmark on hope

What his interns actually did shows where the value went. Instead of accepting the model's first answer, they went to the research literature to work out how you define rubrics for judging agent quality from a trace, and they built a daemon-style collector so the instrumentation cost the developer nothing in latency. The part Chauhan singled out was not the code they wrote. "Those are design decisions, not coding ones," he noted. When the cost of producing code drops toward zero, the scarce input becomes the work a model will not do for you: deciding what to measure, designing the system that measures it, and judging whether the output is any good.

The numbers around him back the diagnosis. In a survey of more than 1,300 teams building with agents, 89% had already put some form of observability in place, but only about 52% ran offline evaluations against a test set. That gap is the diagnosis in one line. Nearly every team can now watch an agent run; barely half have built anything that tells them whether the run was any good. Watching is cheap and nearly universal. Judging whether the run was any good is the scarce part, which is the exact hole Chauhan named when he wrote that "without evals you are not developing agents, you are hoping."

Autonomy raises the stakes

That gap gets more expensive as the agents gain autonomy. Chauhan's instrumentation earns its keep most when no engineer is watching. "Take the human out of the loop and the trace is your only record," he wrote of cloud agents running unattended. What shell commands ran, whether any touched something sensitive, how many tokens went in before the thing converged, or whether it converged at all, none of that is answerable after the fact without the discipline to capture and judge it. The more the machine does on its own, the more the work becomes deciding what good looks like and holding the system to it.

The higher hire

For those running a team, that redraws the hiring signal. If code volume is close to free and every competitor is buying the same agents, raw output stops telling you who the strong engineers are. The people worth paying for are the ones who can define the rubric, read the trace, catch the design flaw the model shipped past, and stand up the evals that separate a working system from a hopeful one. Chauhan called those skills "exactly what all engineers need to evolve rapidly in today's era," and he found them in two interns rather than in a headcount cut. Reward throughput and you'll keep mistaking motion for progress; reward judgment and you'll find you need more strong engineers, not fewer.

Anyone can generate the code. The rare, rising skill is being able to look at what came out and say, with evidence, whether it actually works.

The best candidate is not in your city.

SiiRA connects US companies with top international talent — end to end, effortless.

Hire Top Talent

The best candidate is not in your city.

SiiRA connects US companies with top international talent — end to end, effortless.

See talent differently.

Get the latest ideas on hiring, leadership, and the future of work.

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.