Skip to content

Your 'AI Agent' Is a Pipeline in a Trench Coat (That's a Good Thing)

#ai-coding-agents #claude-code #cursor #github-copilot #agent-architecture #developer-tools

The agent I was proud of, until production ​

I built an agent last year and I was proud of it. It had a planner, tools, and a reasoning loop that decided what to do next, reflected on its own output, and chained steps together. In the demo, it made the room go quiet.

Then it went to production. Slow. Expensive. Failing in ways I couldn't reproduce. Same input on Tuesday, different behavior on Wednesday. When it broke, the cause was three "autonomous decisions" upstream that I didn't control and couldn't see.

So I did the unglamorous thing: I rewrote it as a boring, linear pipeline. Fixed steps. No reasoning loop. It was better on every axis that mattered: faster, cheaper, testable, debuggable.

Then I read the logs from the old agent and felt slightly sick. It had done the same three steps every single run. Extract. Transform. Respond. It never once used its autonomy to do anything different. I had built a for-loop, given it a system prompt, and called it an agent.

I'm not alone. Most of what gets called an "agent" in 2026 is a pipeline in a trench coat. That's not an insult. It's a relief.

One line does most of the work: if you can draw the flowchart of what your system does before it runs, you don't have an agent. You have a pipeline. Step one: retrieve context. Step two: call a tool. Step three: format a response. If you could have drawn that on a whiteboard before writing code, the model isn't deciding the path. You already decided it. The model is narrating your fixed route in fluent natural language and calling the narration "reasoning."

Agency is a cost, and the bill is itemized ​

The distinction that matters: an agent decides its own control flow at runtime. Which tool to call, which step comes next, whether to loop again, when to stop. A pipeline has that control flow fixed at design time. An LLM doing smart work inside a fixed step isn't agency. It's a smart function call.

"Fine," you might say, "so it's technically a pipeline. But it works, so who cares?" Everyone who has to run it, pay for it, or debug it at 2 a.m. The bill is itemized.

  • Nondeterminism. Same input, different path, different output. Bugs stop reproducing. "It worked when I tried it" becomes a permanent state.
  • Debuggability collapse. A pipeline breaks at a step you can name. An agent fails at step 12 because of a decision it made at step 4 that you didn't control and can't replay. Debugging becomes forensics on a choice.
  • Multiplied failure surface. Five fixed steps have five things to check. Five autonomous decisions can each be wrong, in combination, in an order that changes every run.
  • The token bill. A reasoning loop thinks, re-thinks, reflects, and decides to loop again. You pay per token for the model to deliberate about a route you already knew.
  • Untestable. Regression testing needs a fixed set of paths. An agent, by definition, doesn't have one.

Add it up and the punchline is brutal: you paid all of that to let the model decide something you already knew the answer to.

A comment on the original post sharpened the boundary better than the post did. The right test isn't "does the path vary?" It's "is verifying the outcome cheaper than reasoning about it?" A scraper walking a changed DOM, or a retry loop around a 429, earns its runtime freedom because each step's output is cheap to check. Did we get the data? Did the request succeed? My extract-transform-respond loop had no per-step verification, which is why its autonomy bought nothing. Freedom is fine when something checks it. Without a check, it's just a costume.

Quick Take: The difference between a demo agent and a production agent is cheap per-step verification, not smarter prompts.

The model brings the intelligence, and the harness gives it hands ​

There's a second idea hiding in this debate, and the learn-claude-code repository states it better than anything else I've read: agency comes from model training, not from external orchestration code. The historical record is consistent. DeepMind's DQN learned Atari from raw pixels in 2013. OpenAI Five played 42,729 public Dota 2 games at a 99.4% win rate in 2019. AlphaStar reached Grandmaster in StarCraft II the same year. In every case, the intelligence came from gradient updates, not scripted rules.

What most of us actually build is the harness: the environment the model operates in. The repository defines it as tools plus knowledge plus observation plus action interfaces plus permissions.

  • Tools give the agent hands: file I/O, shell, network, database, browser.
  • Knowledge gives it domain expertise, loaded on demand, not upfront.
  • Observation is the feedback: git diffs, error logs, browser state.
  • Action is the CLI commands, API calls, and UI interactions.
  • Permissions are the trust boundaries: sandboxing, approval workflows.

The loop around all of this is almost embarrassingly simple.

The model decides when to call tools and when to stop. The code just executes what the model asks for. Every mechanism that makes an agent product useful sits around this loop: subagents with fresh message lists, context compaction, a persisted task system, permission governance, hooks, memory. Claude Code is the most complete example of this philosophy, and the reason it works is what it doesn't do. It doesn't impose rigid workflows or substitute hand-crafted decision trees for the model's judgment. It gives the model tools, knowledge, context, and boundaries, then gets out of the way.

Even OpenAI now treats coding agents as a research acceleration layer, publishing early data on agent usage, experiment velocity, and task complexity. The same loop that edits a repository is writing training pipelines. At that scale, the question is the same one from earlier: which steps stay fixed, which earn their autonomy, and where the cheap check lives.

The takeaway for builders: you are not writing intelligence. You're building the world that intelligence lives in. The quality of that world determines how well the model can express itself.

A month of pair programming: context beats generation speed ​

I spent the next month running the same test across Cursor, GitHub Copilot, and Claude Code. Code generation speed was never the differentiator. Context was.

The tasks looked like real work in an existing project: add a missing function, debug actual errors, refactor without changing behavior, write tests, ship a feature across multiple files. All three tools produced a working starting point for simple tasks. The gap widened as the scope grew.

Copilot was best at inline autocomplete and small functions, the low-friction stuff you want while typing. Cursor handled interactive refactoring and multi-file changes well, because you can inspect and steer proposals inside the editor. Claude Code was strongest when the task required repository-level investigation before writing any code: trace the data flow, find where the API response stops matching expectations, understand how a function is used across the codebase before touching it.

Programming taskTool that handled it bestWhy
Inline autocompleteGitHub CopilotFast suggestions while typing
Small functions, boilerplateGitHub CopilotLow friction
Interactive refactoringCursorEditor-based review workflow
Multi-file feature changesCursorEasier to guide and verify
Debugging a single-line errorGitHub CopilotQuick contextual fix
Debugging across filesClaude CodeTraces data flow to the root cause
Understanding an unfamiliar repoClaude CodeRepository-level investigation
Writing unit testsAll threeGood start, still needs review
Final code reviewHumanAI shouldn't be the final authority

Debugging was the most revealing test, because the crash site is rarely the cause. Claude Code's ability to follow data across files caught bugs the other two couldn't see without extra context. The testing lesson applied to all three. Prompts that said "write tests for this function" produced tests that blessed whatever the implementation did, including its bugs. Prompts that said "write tests based on the expected behavior, and don't assume the current implementation is correct" produced useful tests. Small change in prompt, large change in what the tests catch.

A practical metric emerged from the month: how much work do you have to do after the AI finishes? Fixing wrong assumptions, removing unnecessary code, rewording error handling, reverting regressions. A tool that generates 200 lines in a minute isn't faster than one that generates 80 useful lines if you spend 30 minutes fixing the first result. Useful code beat generated code every time.

Someone has to supervise: the console layer ​

The most concrete lessons came from a project I started in May: a native console for supervising two competing coding agents, Claude Code and Codex, side by side. The thesis: the agent is the editor, and the human needs a cockpit, not another code buffer. The console's center is deliberately small: a PTY terminal that launches the agent, a git diff view, and an approval and snapshot loop. You don't read code in it. You direct the agent and verify its work.

Designing for two engines from day one was the most profitable constraint. Permission stores have philosophies. When a user clicks "always allow," Claude Code persists that to settings.json, while Codex wants an execpolicy prefix rule in .codex/rules. Same button in the UI, radically different write path underneath. If you integrate only one agent, you'll bake its philosophy into your core and port painfully later. Supporting two forces you to find the real seams.

Convergence is real too. While wiring hooks for both engines, I found that Codex had adopted Claude Code's hook system with schema-identical payloads. Nobody announced it. You only find it by building against both. The flip side: you inherit two release treadmills, and a CLI update can silently restructure its event schema and break your parser. Churn is a permanent line item.

Trust needs evidence, not vibes. Every session in the console is witnessed: prompts, approvals, tool results, and per-turn diffs land in a hash-chained ledger, exportable as signed proof packets. That proves the sequence of events happened in this order and wasn't edited after the fact. It doesn't prove the code is good. It's an audit trail, not an oracle. When a client asks what the AI did to their codebase, "here's a verifiable packet" beats "here's my scrollback."

The war stories mattered more than the features. WebKitGTK's clipboard API on Linux resolves successfully and writes nothing, so the terminal's copy was broken in three different ways at once. Windows hooks failed silently for weeks because a bare script path works on Unix but not on Windows with a space in the home directory. Every fix looked correct in code review; only a human clicking in the real app could confirm them. GUI bugs demand GUI verification.

Key numbers from the field

  • 90%: the share of "agentic" decisions one team demoted back to deterministic steps before its fleet got reliable.
  • 175 pull requests in three months built Agent Console, most of them written by the agents it supervises.
  • 99.4%: OpenAI Five's record across 42,729 public Dota 2 games in 2019, the cleanest proof that agency is trained, not coded.
  • 1.5 seconds: the hook budget a 3 to 4 second embedding load silently blew past on every first call.

Common pitfalls: what trips people up ​

  1. Accepting a large AI refactor without reading the diff. A cleaner-looking implementation isn't safer. AI removes duplication and quietly changes behavior. It "improves" something intentionally written a certain way because another part of the app depends on the oddness. Rule: never accept a large generated refactor unread. Reviewing commits becomes a core skill, not a ritual.

  2. Writing tests that encode the implementation, not the behavior. If the code has a wrong default value, AI-generated tests will bless it. Prompt for expected behavior, edge cases, and failure scenarios. Tell the model the current implementation may be wrong.

  3. Adding autonomy where there's no cheap check. The verification test settles it. A retry loop around flaky infrastructure earns runtime freedom because success is easy to verify. An extract-transform-respond loop without per-step checks pays for freedom it never uses. If a step's output is hard to verify, pin it to a fixed path.

  4. Baking one engine's permission model into your core. Always-allow looks like one concept until you persist it in two places. Claude wants a flat allowlist in settings.json; Codex wants rule files evaluated by the engine. Build the abstraction early, or plan to refactor when the second agent arrives.

  5. Treating silent failure as success. The worst bug in the console project never fired. A knowledge hook injected memory into agent prompts, but the embedding model took 3 to 4 seconds to load on first call against a 1.5 second budget. The injection had never once run in production, and nothing complained. The metrics measured a code path that never executed.

The fix isn't more instrumentation in the abstract. It's instrumenting the success path, not just the failures. A feature that silently does nothing looks identical to a healthy one from the outside, and that asymmetry is the default state of most agent telemetry today.

When the trench coat comes off ​

None of this means real agents never earn their keep. They do, in specific cases. The steps genuinely can't be known in advance: open-ended research, exploration, debugging an unknown system where the path emerges from what you find. Each step depends on discovering the last: real multi-hop work where step 2 is unknowable until step 1 runs. Or the branching is unbounded, not an if-statement with three cases, which is just a pipeline with a switch.

Even then, minimize the agency. Hard-code everything you can. Reserve the model's runtime decisions for the one place that genuinely needs them. Autonomy is a cost. Spend it only where it buys something.

The learn-claude-code lessons back this up with mechanisms. The course states it as a motto: an agent without a plan drifts; listing the steps first doubles completion rate. A later lesson adds an independent evaluator that reviews each proposed stop, so impossible, failed, or over-limit goals return control to the user. That's bounded autonomy: freedom inside a checkable envelope, which is the only kind that survives contact with production.

One thing to remember: the system that ships and stays up in production is almost always more boring than the one that wins the demo. Impressive in a demo was never the goal. Still working on Wednesday was.

The bottom line ​

  • If you're building a production workflow on top of an LLM, start with a fixed pipeline and add autonomy only where a fixed path provably fails and per-step verification is cheap. The teams getting results demoted about 90% of agentic decisions back to deterministic steps; the wins came from that direction, not from smarter prompts.
  • If you're choosing a coding tool for day-to-day work, match it to your workflow: Copilot for inline autocomplete, Cursor for editor-centric multi-file changes, Claude Code for repository-level investigation and cross-file debugging. In every case, keep a human as the final reviewer. The goal isn't zero programming; it's spending your time on the parts that need a human.
  • If you're building tooling on top of coding agents, design engine-neutral from day one, instrument the success path, and budget for vendor churn. Claude and Codex hook schemas are already converging, so a stable agent protocol looks likely within a year, and the durable value sits in the supervision layer: approvals, per-turn snapshots, and verifiable audit trails.