Skip to content

The Bottleneck Moved: What Coding Agents Actually Changed

#ai-coding-agents #software-engineering #spec-driven-development #code-review #verification #developer-productivity

The question stopped being "can it code" ​

Something shifted this year, and it wasn't just the models. The people getting real results from AI stopped using chat and started using agents. Agents that open your codebase, write files, run the compiler, execute tests, and loop on the failures. The fly.io essay made the point hard to ignore: if you're still pasting broken code from a ChatGPT page into your editor, you and the people shipping with these tools are not having the same conversation.

Once you accept that framing, the interesting question changes. Agents can write a lot of code, and they never get tired. So the bottleneck of software engineering moved. It no longer sits in generating code. It sits in everything around it: deciding what to build, deciding what "correct" means, reviewing what the agent produced, and knowing when to stop it from going further.

The evidence from the last year is consistent: agents changed where the hard work lives, not whether hard work exists.

Coding got cheap. Engineering didn't. ​

One story from the community captures the change better than any benchmark. A developer needed to introduce a change that a mediocre coding agent could implement in around five minutes. The work took four days. The delay had nothing to do with code: it was discussing what the change should look like, coordinating with the teams involved, and getting the client to agree on a solution. The decisions changed several times along the way.

Five minutes of coding. Four days of engineering. And an experienced developer in those meetings was absolutely necessary.

This matches the research she points to. METR's studies of experienced developers found that older AI tools sometimes slowed people down. Newer agents do speed things up, but the effect gets murkier the more the task depends on experience, judgment, and coordination. Raw coding benchmarks measure something that is becoming less relevant to the job.

The uncomfortable corollary: if your value as an engineer is mostly "I can turn a specification into code," that skill is now cheap. The part of the job that stays expensive is deciding what the specification should say. One commenter put it well: I use AI to bounce ideas around, but I still make the call. That's not something I want to outsource.

Agent tests can make agents worse ​

Here's a trap I keep seeing. Teams put an agent on a bug fix, the agent writes tests alongside the patch, and everyone assumes green tests mean progress. But an AI-generated test can miss the regression the patch introduces, or worse, encode the wrong expectation into the repair loop.

The ExecCritic preprint from September 2026 measured this directly. Holding the repair agent fixed, they varied only the source of the tests used as feedback:

Feedback sourceTasks resolved
Initial repair, before any generated-test feedback61.2%
Tests from the base Qwen test agent57.3%
Tests from GPT-5.6-sol65.3%

On SWE-bench Verified's 500 tasks, the weak tests cost roughly 20 resolved tasks compared to no feedback at all. Better tests gained about the same amount in the other direction. Tests aren't a neutral instrument. They steer the agent.

Key numbers from the ExecCritic preprint61.2% tasks resolved without generated-test feedback 57.3% with tests from the weaker agent, a 3.9-point regression 65.3% with tests from the stronger model, a 4.1-point gain

The industry is already acting on this. Cognition is pairing Devin with GPT-6 Astra specifically to improve how Devin tests its own work, with the goal of cutting human review load. The ExecCritic numbers suggest that bet is sound, and that it cuts both ways: worse tests don't just miss bugs, they actively push the repair loop in the wrong direction.

The mechanism is easy to demonstrate with a tiny example. A filter function has three requirements: omit the filter or pass None and return everything, pass an empty list and return nothing, pass a list of statuses and return only matches. A plausible patch replaces the None check with a truthiness check:

python
if not statuses:
    return list(orders)

Both obvious checks pass. Both branches get exercised. The reported bug, omitting the filter returns nothing, is fixed. Add the empty-list assertion and the patch fails. Python treats None and [] as equally falsey, but the requirements give them different meanings.

The right question to ask of any AI-generated test is not "does it pass?" but "which plausible wrong implementation would this test reject?" When I ran this fixture against eight wrong implementations, one per contract axis, only an identity check separated list(orders) from return orders, a one-click diff that nothing else in the suite observed. Even a 100% mutation score proved nothing: the same five mutants were all killed for the broken patch and the correct one. The mutation operators were drawn from the code, not from the requirements.

That is why the "fails before the fix, passes afterward" check, the gold standard of test-driven repair, is not enough. It detects the original bug. It says nothing about the next one.

The contract discovery problem ​

The deeper issue showed up in a password reset flow a coding agent built. The reset link worked more than once. The bug survived because nobody had written down that a reset link should be single use. It was obvious right up until it wasn't.

The fix, adding an independently written behavioral spec, worked. The agent implemented against it. The verifier rejected the reusable token. The agent fixed it. Then the comments found holes the spec didn't cover. What if two reset requests with the same token arrive at the same time? Validation and consumption weren't atomic, so "single use" still produced two successful resets. The invariant was incomplete.

That loop is messier than the tidy "spec then implement" diagram. It also happens to be engineering.

Quick take: The hardest problem isn't verification. It's contract discovery: deciding which invariants are worth verifying in the first place.

The nastiest part: separate artifacts don't mean independent verification. One commenter described an integration builder where the agent wrote both the connector and the tests for it. Everything passed. Both were wrong about OAuth token refresh. The implementation and its test suite were separate files but shared one incorrect assumption, same as two students who studied from the same wrong answer key.

So the contract can be wrong. The verifier can be wrong. In one verification harness someone described, a capability test timed out and the harness recorded FAILED when the honest result was NOT TESTED. Deterministic does not mean correct.

This is why the open question, who writes the contract, is harder than it looks. If humans must write complete specs before agents can do anything, the bottleneck just moved back to humans, and the concurrency example showed that humans don't know the spec upfront either.

An agent can propose the contract. That sounds circular, but it isn't, as long as proposing and accepting are different operations. The review target shrinks from four hundred lines of implementation to four lines of stated invariant. Tools like GitHub's spec-kit formalize this loop: a one-time constitution per project, a specify step that captures the what and why, a plan, task breakdown, implementation, and a converge step that assesses the codebase against the spec and appends the remaining work. The point isn't the ceremony. It's that the spec, not the code, becomes the artifact the agent works against and the reviewer evaluates.

The catch: acceptance has to introduce something the proposing agent didn't have. An agent that doesn't know single use matters will not propose single use as an invariant. The value comes from the independence of the source, not from the ceremony of the review.

Nobody installed the stop sign ​

There's a second failure mode, and it's the one that surprised me most. Agents don't get tired, so work doesn't stop. The constraint that used to end a task, running out of hours, is gone.

One developer's story is painfully familiar. They tagged a folder [THROWAWAY] and budgeted two hours for a quick deploy test. A build day later, the throwaway had a byte-identity test protecting a vendored logging module, exact dependency pinning, and a sync script with drift detection. The agent proposed them, a reviewer agent approved them, and the human approved all of it. Every decision was defensible on its own. Nothing had decided whether that thoroughness belonged in a folder marked for deletion.

Satisficing, Herbert Simon's term, is an adaptation to scarcity. You stop when you hit diminishing returns because you have somewhere else to be. Agents have no scarcity. The stop sign has to be installed by hand.

The mechanisms that work are mechanical, not motivational. Rigor tiers work when they're written as prohibitions: "no test files, no dependency pinning, no sync scripts" instead of "spike rigor." Even better are hard walls. One person described git hooks that reject any commit adding a dependency to a spike directory, with the agent run capped at three tool turns. A rule in a prompt is a suggestion. An exit code is a wall.

I found the same pattern in my own work. I prompted an agent to fix a tiny CSS issue and watched it start searching the entire codebase, burning tokens. I stopped it after a few minutes and made the change myself. Arbitration, deciding whether the agent or the human should do a given thing, is now a core engineering skill.

What the community is saying: in the threads I read this year, the loudest debates aren't about whether agents can code. They're about what happens after. Developers review a dozen PRs before lunch from asynchronous agents. They watch agents burn tokens on a five-minute CSS fix while missing a one-line credentials change. They point at the same regex on the same line of the same diff, where nobody can tell whether the reviewer understood it or accepted it. The pattern across all of it: the code was never the hard part. Knowing whether the code is right, and whether it should exist, is.

The kernel says it best ​

The strongest production evidence comes from the Linux kernel. Arm engineer Lorenzo Stoakes used an LLM to hunt build bottlenecks. The AI found real problems in the build pipeline, then participated in building, testing, and debugging. Then it wrote fixes. Stoakes' verdict on the generated code was direct: a lot of it was hideous. He rewrote much of it by hand.

The final result was 23 patches, each tagged Assisted-by, and the build-time gains were large:

For a kernel developer running builds dozens of times a day, that's not a nice-to-have. A few minutes saved per cycle compounds into hours.

The part that matters is the division of labor. The AI enumerated possibilities and put them on the table. A human decided which survived. Linus Torvalds' debugging of the Intel Xe driver follows the same shape: the AI repeatedly claimed the problem was unsolvable and suggested filing a report. The human kept pushing it to add debug code and analyze results. After 24 debug patches and 18 reboots, the answer was a one-line fix, a round_down() that should have been round_up(). The AI did the grunt work. The human supplied the refusal to quit.

Picking libraries for the agent, not the human ​

Library selection criteria are shifting too. At a conference, Joel Hooks showed Effect, which advertises itself as reliable TypeScript for the AI era. His claim: agents write better TypeScript with Effect than he could by hand.

Effect routes expected failures through a typed error channel. Add a new error case, and an exhaustiveness check can flag every UI branch that forgot to handle it. StyleX does something similar for styling, putting visual states behind a typed API so an agent can't invent selectors or misspell properties. The compiler catches errors first.

That raises a strange question. Do you adopt a library because your agent writes it better, even when it's harder for you to learn and maintain?

ChoiceHelps the agentCosts the human
EffectTyped failures, exhaustiveness checksNew programming model, steep learning curve
StyleXNamed visual states, constrained compositionUnfamiliar syntax, generated class names
Plain TypeScriptFamiliar syntax, broad ecosystemFailures stay implicit, agent guesses more

I'm skeptical of optimizing for the agent alone. Exhaustiveness checks only fire in consumers written exhaustively. If the agent writes if (error._tag === ...) checks, a new error lands silently in the else branch. The benefit is bounded by a convention you have to enforce by hand, which is the same kind of tribal knowledge you were trying to replace.

There's also version drift. Agents generate code confidently from training data, and a major version bump can silently break the patterns they reach for. When I tried Effect, I kept producing patterns a version behind. The question isn't just "can the agent write it well now," but "can the agent stay current with it."

The middle ground is something like shadcn/ui: common patterns, source stays in your project, and a human can still open the file and understand what the agent generated. Give the agent enough structure to catch mistakes, but keep the code review legible for the owner.

Common pitfalls ​

  • Treating agent-written tests as trustworthy. A test that passes against the patch and fails against the original only proves the reported symptom is addressed. It says nothing about the regression you just introduced. Review expected values against the requirement, not against the current output.
  • Letting the patch run its own tests. If the agent can modify both the implementation and the test command, a green run is meaningless. Keep a separate evaluator with reviewed tests and a frozen test command, and don't let the patch skip execution.
  • No stop criteria for agent work. Agents over-harden throwaway code because nothing tells them to stop. Write rigor tiers as prohibitions in the prompt, and back them with mechanical walls like git hooks that reject scope creep.
  • Assuming separate artifacts means independent verification. An agent that writes the connector and its tests can encode the same wrong assumption in both. Independence comes from the source of the definition of correctness, not from the pipeline stage.
  • Accepting agent-proposed contracts without challenge. A confident, well-formatted contract with the same hole as the implementation is the worst outcome, because now the hole has been written down and approved. Acceptance must introduce something the proposing agent didn't have: a human who has debugged that class of bug, or a checklist from past incidents.

One thing to remember ​

The implementation is disposable. The accumulated definition of correctness is not. The reset flow may be rewritten next month. The framework will change. The agent will change. But once production teaches you that two competing reset attempts cannot both succeed, that invariant should survive all of those changes. Write it down somewhere versioned, reviewable, and attached to the behavior rather than to the incident that revealed it.

The bottom line ​

  • If you're running agents on a real codebase, adopt a spec-driven workflow with the spec written before the agent starts. Reviewing four lines of contract is a conversation about intent. Auditing four hundred lines of code is an audit, and you won't do it.
  • If your agent loop keeps producing regressions that look like test failures, add an independent evaluator. Freeze the reviewed tests, keep the test command out of the patch's reach, and run one deliberate wrong implementation to confirm the suite would reject it.
  • Watch the move from agents writing code to agents calling APIs. A bad tool call already happened: the email was sent, the card was charged, the access was revoked. There's no diff to read afterward. Defining what agents are allowed to cause, and what evidence proves the intended effect happened, will be the defining engineering problem of the next year.