Skip to content

Agentic Workflows Are Going Production. The Specs Don't Travel.

#ai-agents #agentic-workflows #mcp #sre #deep-research #llm-tools

Agentic Workflows Are Going Production. The Specs Don't Travel. ​

Agents got real jobs this year. One framework ran a job search that ended in a signed contract. Another caught a four-day ECS failure loop while its owner slept. A third backtests trading strategies with 286 tested quantlib functions. These aren't demos. They're workflows with tool access, evaluation loops, and safety boundaries.

The connective tissue is MCP. Kiro Crew adds AWS DevOps Agent with one config block. Vibe-Trading exposes 74 MCP tools. OpenSRE connects 60+ tools across observability, infrastructure, and incident management. open_deep_research has full MCP compatibility. The pattern is consistent: agents are only as useful as the tools they can reach, and MCP is how they reach them.

But there's a catch, and it's hiding in a new arXiv paper. The artifacts these agents produce don't travel between agents. When a specification written for one agent gets handed to another, you can end up with a 2.33% SQL syntax validity rate. The transfer doesn't degrade gracefully. It falls off a cliff.

The portability problem, measured ​

The paper, "Specification Portability Across LLM Development Agents," uses Oracle-to-PostgreSQL migration as a controlled transformation task. Stage one: a specification-first migration pipeline on 1,006 PL/SQL files. Stage two: cross-agent experiments on 1,802 Oracle scripts using Amazon Kiro, Google Gemini, and GitHub Copilot, with Claude Code and Cursor in the initial single-agent evaluation.

The single-agent baseline tells you how hard this task is even without the cross-agent wrinkle:

StageFilesShare
Input PL/SQL files1,006100%
Regenerated by spec-first pipeline62362%
Executed successfully in PostgreSQL 1638038%

Even in the best case, a spec-first migration loses 38% of files at regeneration and another 24% at execution. That's the baseline. Now add the cross-agent transfer.

Key Numbers

  • 623 of 1,006 PL/SQL files survived the spec-first pipeline (62%)
  • 380 of 1,006 executed in PostgreSQL 16 (38% end-to-end)
  • 0.035 Token F1 when Gemini directly consumed a Kiro-origin spec
  • 2.33% SQL syntax validity in that same transfer

The worst replicated case: Gemini directly consuming a Kiro-origin specification produced a Token F1 of 0.035, SQL syntax validity of 2.33%, and AST mean similarity of 0.015. Token F1 of 0.035 means essentially zero token-level overlap with the reference implementation. A 2.33% syntax validity means 97.7% of the generated scripts don't even parse. The output is effectively a different language.

Why specs don't travel ​

The paper tested three mitigation strategies: rewriting the spec, compressing it, and retrieval-augmented ingestion. The findings are uncomfortable for anyone building multi-agent pipelines.

Specification size alone doesn't predict implementation quality. Rewriting substantially improved Gemini in the tested configuration. Compression provided no universal benefit. And retrieval-augmented ingestion was the only strategy that landed on the per-agent Pareto frontiers of both Gemini and Copilot.

The practical read: a specification is not a neutral artifact. It's an encoded conversation between a human and a specific agent's priors. When you hand that spec to a different agent, you're handing it a document written in someone else's shorthand. The agent doesn't know what it doesn't know, so it fills the gaps with its own defaults, and those defaults are wrong often enough to matter.

Quick Take: A specification is an encoded conversation between a human and a specific agent, so cross-agent transfer without retrieval-based access isn't degraded performance, it's a different language.

The capability layer pattern ​

Agent Reach attacks a different wall: agents can write code, but they can't read Twitter, search Reddit, or pull YouTube subtitles. Each platform has its own barrier. Paid APIs. Anti-bot blocks. Login requirements. Data that comes back as a pile of HTML tags.

The project calls itself a capability layer, and that's the right word. It doesn't read anything itself. It selects, installs, health-checks, and routes. Each channel file probes candidate backends in order and picks the first fully working one.

The routing isn't theoretical. I had yt-dlp die on me when Bilibili's anti-bot started returning 412s in June 2026. The fix was reordering one channel file to prefer bili-cli. No code rewrite, no user action. That's the entire point of the layer: backends get replaced, the interface stays.

The same abstraction shows up in every serious framework in this cluster. Vibe-Trading's IBKR connector went from "lists tools" to a working read-only portfolio source through a series of OAuth fixes, and the agent sees exactly one scheduling tool whose create and cancel calls never touch the job store until a human confirms. OpenSRE masks sensitive identifiers before external LLM calls and restores them in output. The pattern is consistent: abstract the flaky upstream, keep a fallback, and never let the agent's reach exceed its authority.

Workflow as the product ​

The ai-job-search project is the most honest agent application I've seen this year. The author is a geophysicist whose position was cut in late 2025. He built a Claude Code framework to run his own job search, used it weekly, and was upfront with every employer about it.

I ran this on my own job search. Sixty-nine tailored applications, twenty first interviews, one signed contract. I started as an AI engineer in June 2026. That's roughly a 29% interview rate, which most job seekers would take in a heartbeat. The author notes that disclosing the agent usually sparked a genuine technical conversation instead of counting against him.

The core pattern is a drafter-reviewer pipeline. /apply parses the posting, evaluates fit against a structured scoring framework, drafts a CV and cover letter in LaTeX, spawns a second agent to research the company and critique the drafts, revises, compiles both PDFs, and runs an ATS parseability check. Postings are treated as untrusted input: the workflow follows no instructions embedded in them and fetches no links from their body.

The interview prep command is where the design gets interesting. It builds a stage-specific pack from the application archive: the exact posting, the CV and cover letter the interviewer actually read, feedback recorded from earlier rounds. It maps likely questions to STAR examples and runs a mock interview. Gaps get honest bridge answers, never invented experience. That constraint, no fabrication, is what makes the whole thing defensible to an employer.

Domain agents: SRE and trading ​

The Kiro Crew story is the one that made me set up my own cron-based checks. At 3:17 AM, a payment-api ECS service entered a terminal failure loop. 7,279 failed tasks since August 13th, burning compute the whole time. The health check expected /api/health but the container only served static content. Every 60 seconds, ECS killed the task and replaced it. No alarm fired because none was configured. The author slept through it and woke up to a solved problem: five parallel investigations, root cause found across ECS, CodeBuild, CodePipeline, and Lambda, severity-prioritized fixes flagged. Done by 3:24 AM.

The architectural insight is the brain/hands split. AWS DevOps Agent is read-only by design. It observes, correlates, and produces mitigation plans with exact commands, but it never executes. Kiro Crew is the hands: it orchestrates, executes, verifies, and learns from past incidents. In production, fixes become pull requests for human review. Neither half alone solves the problem.

OpenSRE pushes the same philosophy one level deeper. It's not just an SRE agent, it's a training and evaluation environment: scored synthetic RCA suites that check root-cause accuracy, required evidence, and adversarial red herrings, plus real-world end-to-end tests across Kubernetes, EC2, CloudWatch, Lambda, ECS Fargate, and Flink. The mission statement is sharp. SWE-bench gave coding agents scalable training data and clear feedback. Production incident response still lacks an equivalent. OpenSRE wants to be that equivalent. It's public alpha, and the benchmark table is still empty. That's the honest part: the environment exists before the results.

Vibe-Trading shows what correctness obsession looks like in a domain where it matters. The changelog reads like a confession log: a NaN peer metric dragged the median to NaN; abs(nan) > tolerance is False, so a NaN balance sheet sailed through the hard balance check; search_symbol returned zero candidates with both sources reporting ok. The fixes are the instructive part: refuse non-finite inputs everywhere they enter the arithmetic, fail closed on stale data, and keep the portfolio path read-only so nothing on it can place an order.

FrameworkDomainCore patternSafety boundary
ai-job-searchJob huntingDrafter-reviewer pipelinePostings untrusted; human sends
Agent ReachWeb accessCapability layer + fallback routingCredentials local (file mode 600); dry-run install
Kiro Crew + AWS DevOps AgentSRERead-only brain, executing handsPR approval in production
OpenSREIncident responseEval-driven RCA agentsIdentifier masking before external calls
Vibe-TradingTradingRead-only connectors + 74 MCP toolsPortfolio path can't place orders
open_deep_researchResearchPlan-execute with parallel searchConfigurable models and search per task

The cost of autonomy ​

Deep research is the reference application for agentic workflows, and open_deep_research is the open reference implementation. It hit #6 on the Deep Research Bench leaderboard with a 0.4344 RACE score. The benchmark is 100 PhD-level research tasks across 22 fields, judged by Gemini against expert-compiled golden reports.

The cost table tells you more than the scores do:

ConfigResearch modelCostTokensRACE
Defaultsgpt-4.1$45.9858.0M0.4309
DRB submissiongpt-4.1$87.83207.0M0.4344
Claude Sonnet 4claude-sonnet-4$187.09138.9M0.4401
GPT-5gpt-5n/a204.6M0.4943

The spread between 0.4309 and 0.4401 is small. The cost spread is 4x. Claude Sonnet 4 costs $187 for the same benchmark the default config runs for $46. If you're running research agents in production, model choice is a budget decision, not just a quality decision. A 0.01 RACE improvement for 4x cost is hard to justify unless you're chasing a leaderboard slot.

Common Pitfalls ​

  1. Treating specs as agent-neutral. This is the paper's core finding, and it applies beyond database migrations. If you're handing off prompts, specs, or generated artifacts between agents, test the transfer on a sample before committing. Use retrieval-augmented ingestion rather than stuffing the spec into the prompt.

  2. Trusting ok flags. Vibe-Trading's changelog is a graveyard of tools that reported success while failing. search_symbol returned zero candidates with both sources reporting ok. get_options_chain answered a wrong-cycle expiration with ok: true and another date's contracts. Even the project's own test suite was appending fabricated order_rejected records to the live hash-chained audit ledger until someone sandboxed the config root. Build assertions that check outputs against ground truth, not against the tool's self-report.

  3. Running agents without basic observability. The Kiro Crew story only worked because the agent checked every 30 minutes. The author admits no alarms were configured. The agent caught the failure, but the lesson cuts both ways: autonomous agents complement alarms, they don't replace them. The agent caught it faster because it was watching, not because the monitoring was good.

  4. Letting the same agent investigate and execute. The strongest pattern in this cluster is the separation. AWS DevOps Agent is read-only. Kiro Crew executes but opens PRs in production. Vibe-Trading's portfolio path can't place orders. ai-job-search treats postings as untrusted input. When one agent both diagnoses and acts, you lose the audit trail that makes the diagnosis trustworthy.

  5. Ignoring the cost curve on open-ended tasks. Deep research runs burn 58M to 207M tokens per benchmark pass. That's $46 to $187 per 100 tasks. In production, set a token budget and a max-iteration cap, or a single research request can outspend its value.

One Thing to Remember ​

Every framework in this cluster that works has a hard boundary between the agent's judgment and the agent's reach. The brain can investigate, draft, and recommend. The hands need a human in the loop or a read-only constraint. The moment you blur that line is the moment a prompt injection or a bad tool output becomes an incident instead of a draft.

The Bottom Line ​

If you're building multi-agent workflows where agents hand off specs or prompts to each other, don't assume portability. Test the transfer on a sample and use retrieval-augmented ingestion, because the paper shows it's the only strategy that lands on the Pareto frontier for multiple agents.

If you're running agents against production systems, copy the brain/hands split: read-only investigation plus execution that requires PR approval. Kiro Crew and OpenSRE both converge on this design, and it's what makes the 5-7 minute MTTR story safe to replicate.

If you're deploying agents for open-ended tasks like research, watch the cost curve. The 0.01 RACE gap between the $46 and $187 configs isn't worth 4x spend unless you're on a leaderboard. Set token budgets before you set expectations.