Appearance
The failure that doesn't announce itself
Two months after wiring an AI agent into social platforms, I learned that the plumbing was never the risky part. The OAuth flows, the token refresh schedules, the media upload pipelines, all of that got solved and hidden behind a single tool call. What I missed was simpler and scarier: once an agent has write access, failures stop announcing themselves.
Ask an agent to post to LinkedIn and it can hand the API a value with the right prefix, the right length, the right shape, and the wrong account. Nothing about that request is malformed. There is no error to catch. The API was asked to do something specific and it did it. The comment thread gave it a name: semantic authorization. The platform validates that the identifier is structurally valid and still accepts an action against the wrong target. The system reports success, and the user gets something they didn't ask for.
That's the failure class this field is now designing against. The multi-agent systems being shipped today don't crash often. They succeed at the wrong thing, quietly, and you find out later.
The research from the last few months is converging on the same conclusion from different directions. ProgRouter attacks the cost of orchestration. SymTrace attacks the assumption that rerunning a failed trajectory is debugging. OpsHarness attacks the instinct to build a specialized agent from scratch. Warp attacks the fact that feedback disappears when the session ends. None of them make the model smarter. All of them make the system around the model more reliable.
Routing is a step-by-step problem
Multi-agent workflows burn money in two ways: repeated LLM invocations and long-horizon context accumulation. Every step that routes to the wrong model is wasted tokens, and every wasted step extends the context that makes the next steps more expensive.
Existing cascade routing methods treat this as a one-shot decision. Pick the right model for the query, done. ProgRouter makes the case that this is wrong for multi-step workflows, because the right LLM at step four depends on how the task has progressed by step four. The problem is state-dependent, so the routing decision has to be made online, per step.
ProgRouter's design has three pieces: a multi-view task progress scorer that combines coarse outcome regimes with fine-grained signals on subtask completion and progress trends, a dual-path predictor that estimates the progress gain for each candidate routed LLM, and an adaptive meta-gating mechanism that balances progress gain against time budget and long-term cost efficiency. The routing decision at each step is: does paying for the expensive model now buy enough progress to justify the cost?
The results on HumanEval Plus, MBPP, MATH-500, and ASQA show the approach holds quality while cutting operating cost relative to key baselines, across agentic code generation, mathematical reasoning, and retrieval-augmented long-form QA. The practical implication: you don't need a cheaper model. You need a cheaper decision about which model to call.
Quick Take: the field has stopped trying to make agents smarter and started trying to make the systems around them trustworthy.
Rerunning is not repairing
The uncomfortable question at the center of multi-agent debugging: when a rerun fixes a failure, did the system actually repair anything, or did it just win the sampling lottery?
SymTrace, a controlled evaluation framework from a recent paper, answers this with data. It records execution trajectories, establishes intervention anchors, replays the execution before the anchor from logs, and only regenerates the downstream trajectory. That lets researchers reproduce failures reliably instead of hoping they recur. The team built SymFail, a dataset of 536 human-annotated failure trajectories with graph-linked locations, categories, and trace evidence, enough failure cases that the patterns aren't anecdote.
The findings are brutal. Unguided reruns reproduce the failure only 67.97% of the time, and repair it only 6.90% of the time. Roughly a third of the time, the rerun doesn't even hit the same failure, and when it does, it fixes it less than one time in ten. Most existing MAS debugging methods are not causally repairing failures. They are stochastically repairing by leaning on the randomness of LLM sampling.
The result that matters more: a symptom-driven intervention method, which uses the trace evidence to guide regeneration at the failure point, repairs 20.15% of failed cases. That's a 191.89% improvement over existing repair methods. It's still a low absolute number, which tells you how far the field has to go, but it's evidence that guided repair beats blind resampling.
Build the harness, not the agent
OpsHarness starts from a quantitative observation that runs against the grain of a lot of agent engineering: general-purpose agents like Codex or Claude Code now beat specialized RCA agents built from scratch for root cause analysis. The accuracy of the general agent still falls short of production needs, but the gap comes from the external adaptation layer, the harness, not from the agent's general capabilities.
The paper's argument: stop rebuilding the agent and start building the harness around it. A key capability of that harness is self-evolution, accumulating system-specific experience from past diagnoses so it gets better the more it's used.
The OpsHarness architecture splits into a data plane and a control plane. The data plane combines layered operational knowledge with an idea-card tool library. The control plane coordinates setup, diagnosis, evolution, and verification. During evolution, the system contrasts successful and failed trajectories, converts their evidence into atomic proposals, and only admits updates through a dual-gate verification process designed to prevent overfitting and regression.
Key numbers
- 59.0% top-1 accuracy on RCA benchmarks, a 63.4% improvement over a bare general agent.
- 4.02x the accuracy of baseline RCA agents built from scratch.
- Measured across two public benchmarks and an industrial deployment, so the numbers aren't lab-only.
That 59.0% is the difference between a tool that points at the right service and one that sends an on-call engineer on a tour of the wrong ones. The dual-gate verification matters more than it sounds: it's the mechanism that keeps the harness from learning the wrong lessons from a single noisy incident.
The same harness philosophy shows up in unexpected places. OpenMontage, an open-source agentic video production system, is pipeline-driven: the agent guide tells agents to read the contract first and not improvise the production workflow. Tool discovery goes through a registry, stage directors live in skills, and a local board shows what the pipeline is actually doing while a run is in progress. Scene-by-scene approval gates pause asset generation until a human signs off on the visuals, and a replay mode scrubs through a whole production from its timestamps. The cost numbers grab attention, a 60-second animated short for $1.33 in API fees, but the orchestration design is the part to study even if you never touch video.
Feedback that compounds
Warp's self-improving agent pattern is the simplest thing in this group, and maybe the most transferable. The problem: feedback to an agent disappears when the session ends. An engineer thumbs down a bad code review comment, the agent moves on, and the next run starts from the same ignorance.
Warp's fix is two skills and a human in between. The inner skill holds the functional domain knowledge for the task. The outer improver skill runs on a schedule, pulls accumulated human feedback, compares what the agent suggested against how humans responded, and proposes a small, focused edit to the inner skill. Because skills are plain files, agents are good at updating them, and the updates flow through a normal PR and code-review workflow. Merge the PR, and the next run of the inner skill inherits the improvement.
The details matter. Warp's issue triage agent missed a label on a sample issue, a maintainer left feedback directly on the issue explaining both what he expected and why, and the improver skill turned that into the smallest edit that captured the signal. The bundled Python script that pulls recent issues and summarizes them into JSON is a proven pattern: skills can reference resource files instead of writing fresh code on every run.
This pattern works because it keeps a human in the loop at exactly the right point: the merge. The agent proposes, the human disposes, and the improvement compounds across the organization. Warp runs it across its entire open-source repo with separate spec-writing, review, and triage agents, each carrying its own loop.
Evidence is the real bottleneck
The deepest failure mode in agentic systems isn't routing or debugging. It's that the system's own records can be confidently wrong.
I had a gate that decided whether an agent could act on a permission grant. When it refused, it wrote a row explaining why. One field held the reason the conditions had changed, and I was storing a label there, my own assertion about what happened. A reviewer told me to store the raw before and after instead, so a stranger could recompute the verdict without believing me. I applied that constraint where the comparison happens, and then found out I had only obeyed it in one direction. One file downstream, a classifier was reading an English sentence in the notes field to decide what kind of evidence it was, grepping for the words "ttl expired" instead of checking the structured TTL number sitting on the same event. Rename the note and the evidence classification changes.
The fix looked obvious: a typed enum. Then the reviewer asked one question: what proves the enum was true? If the classifier trusts the enum without checking the structured TTL, you haven't removed a self-assertion, you've retyped one. The row he wanted to try had source_consult = SKIPPED_TTL_EXPIRED next to ttl_remaining_hours = +17.4. The evidence class said the grant expired. The number said it had seventeen hours left. It returned TTL_EXPIRED.
Then the rounding bug. The gate stored ttl_remaining_hours = round(seconds / 3600, 2). A grant one second past expiry stores -0.0, and in Python, -0.0 >= 0 is True. Every grant expired by less than about eighteen seconds was misread. Not because the clock was wrong, but because the classifier was making an evidence-class decision from a rounded display copy of the clock while the exact timestamps sat on the same object. And nan fell straight through the guard: float("nan") >= 0 is False, so malformed evidence produced a confident answer.
The lesson, stated in the article that documents this: structure is not evidence merely because it has a schema. A typed field can lie as cleanly as a sentence. The fix moved expiry authority to the direct comparison decision_timestamp > grant_expires_at, computed without rounding.
This is the same disease that shows up in the reasoning ledger design conversation: derived values outranking the raw evidence on the same row. Steal these principles wholesale. The ledger witnesses, it does not enforce. Supersession is a new event, never a rewrite. Record how the authority was obtained, not just which one. Relationships need two clocks, valid time and transaction time, or you'll judge a past decision using knowledge that arrived in the future. Preserve what lost, not just what won, because a ledger that records only the supporting evidence is a post-hoc justification engine wearing an audit trail.
The comment thread on the ledger post sharpened this further. One commenter put it in a way I haven't stopped thinking about: success that nobody independent checked is not success, it is only a report. I've had a crash that looked like database failover, and every signal agreed, because all of them measured the same wrong thing. That's the multi-agent version of the problem: the agent's confirmation is not an independent check, because it's the same model agreeing with itself.
The table I keep coming back to, comparing the four systems in this group:
| System | Problem it targets | Core mechanism | Headline result |
|---|---|---|---|
| ProgRouter | Cost of multi-step orchestration | Online progress-guided routing with multi-view task progress scoring | Lower operating cost at maintained quality across code, math, and long-form QA |
| SymTrace | Failure debugging | Controlled replay with intervention anchors; symptom-driven regeneration | 20.15% repair rate vs. 6.90% for unguided rerun |
| OpsHarness | Root cause analysis | Self-evolving harness around a general agent; dual-gate verification | 59.0% top-1 accuracy, 4.02x over baseline RCA agents |
| Warp skills loop | Recurring task quality | Inner and improver skill pair with human feedback; PR-based updates | Feedback compounds across runs instead of vanishing |
Different problems, same shape: the model stays put, and the reliability comes from the system around it.
Common pitfalls
Don't trust a green response as proof of a correct side effect. The agent posted successfully, to the wrong account. The request was structurally valid, the API returned success, and the failure was invisible until it mattered. Resolve the target identity server-side immediately before any write, echo it back, and log the target account next to the action ID. Treat the expected target as a separate invariant, not as metadata carried through the agent's context.
Don't rerun failed trajectories and call it debugging. A rerun that fixes a failure is not evidence of a repair. SymTrace's numbers say you'll reproduce the failure about two-thirds of the time and fix it about seven percent of the time. Capture the trace, anchor the failure point, and regenerate only the downstream trajectory, guided by the evidence. Blind resampling is a lottery ticket.
Don't store derived labels instead of raw evidence. The enum that replaced the prose note was an improvement, and it was still wrong, because it could contradict the raw field sitting on the same row. Store the least-derived evidence available, and make the authority path replayable from the row alone. If a stranger can't recompute your verdict from the row, your record is an assertion wearing a computation's clothes.
Don't build a specialized agent from scratch when a harness would do. OpsHarness's quantitative finding is that general agents now beat specialized RCA agents, and the gap is the harness. If you're building an agent for a narrow domain, start from a strong general agent and invest in the adaptation layer, the tool library, and the experience accumulation, not in re-deriving agent capabilities.
Don't treat passing tests as proof the spec is right. Twice in that permission-gate story, the implementation was faithful and the specification was wrong. 366 tests passed against a contract that mandated a contradiction. Passing tests measure conformance to a document, not whether the document is right. The cheapest way to find that gap: describe your contract in plain sentences to somebody who can't run it, and let them tell you what your own words permit.
One thing to remember: the expensive failure in a multi-agent system is the one that returns success. Exceptions land in logs and somebody eventually reads them. The call that succeeds and quietly does the wrong thing, right shape, wrong target, no error anywhere in the chain, that's the one that gets published.
The bottom line
If you're running agents with write access to anything outward-facing, build a dry-run target and server-side identity resolution first. The playground endpoint that validates against real rules and throws the result away turned out to be the most valuable guardrail in the whole system, because it made the entire round trip testable without an audience.
If you're debugging a multi-agent system that fails on long-horizon tasks, stop rerunning the whole trajectory and start capturing trace evidence with intervention anchors. Symptom-driven regeneration at the failure point repairs three times as many cases as blind resampling, and it tells you something about why the failure happened instead of just whether the lottery paid out.
If you're building agents that handle recurring tasks, adopt the two-skill loop: a base skill that holds domain knowledge and an improver skill that turns human feedback into reviewable, mergeable edits. It's the only pattern in this group that gets better with use, and it costs almost nothing to bolt onto an existing agent. One thing to watch: the reasoning ledger conversation is heading toward write-side custody, tamper-evident records rather than just well-designed ones. That's the next layer of the reliability stack, and it's coming faster than most teams are planning for.