Appearance
Four systems, one pattern
Agent research is converging on an awkward conclusion: after a year of single-agent hype, the strongest results are coming from structure, not scale. Not bigger models, but deliberate architecture. This week's papers and releases all make the same move, away from flat single-turn LLM calls and toward organized systems.
The move shows up at every layer of the stack:
| System | Core idea | Reported result | What it means in practice |
|---|---|---|---|
| Trace2Tower | Transition-aware skill hierarchy induced from execution traces | 87.31% success on ALFWorld, 50.67% exact match on WebShop | The agent finishes roughly 9 of 10 household tasks, and half of shopping runs end with the exact correct item |
| First Things First RL | RL over must-have vs. nice-to-have requirements | Large gains over strong baselines on 3,649 problems | Models stop violating hard constraints when priorities are explicit and abstention is allowed |
| DMoA | Structured role-based debate between agents | +10.21 pp accuracy, +11.36 pp safety over GPT-4o | On complex clinical cases, the debate format beats a single model by a margin that matters for diagnosis |
| DeerFlow | Open-source agent harness: sub-agents, memory, sandboxes | #1 on GitHub Trending after the 2.0 rewrite | A production-grade agent stack you can run locally, with an actual ops story |
| AutoHedge | Swarm of specialized trading agents with a risk-first pipeline | Autonomous end-to-end trading on Solana | Agents generate theses, size risk, and execute, with risk gating between analysis and action |
Read them together and the lesson is consistent: agent quality is becoming an architecture problem.
Key numbers
- 87.31% success on ALFWorld, in 10.35 steps with 0.26 invalid actions per task: the agent finishes nearly 9 of 10 tasks and wastes almost no context on dead ends.
- 50.67% exact match on WebShop: half of shopping sessions end with the precisely correct item under a strict metric.
- +10.21 pp diagnostic accuracy and +11.36 pp safety over GPT-4o: the debate structure shifts clinical outcomes, not just benchmark scores.
- 3,649 problems in FTF-rl: enough coverage of e-commerce, booking, map-based, and ride-hailing scenarios for the failure pattern to be credible.
Trace2Tower: raw traces become a skill tower
Execution traces are rich, but most systems treat them as a flat pile of text. You retrieve a similar trajectory, stuff it into the prompt, and hope. Trace2Tower's authors argue this ignores the temporal dependencies and outcome-conditioned topology of agent behavior, and they're right. A trace isn't a document. It's a path through a state space where some steps lead to success and others to loops.
The pipeline works in layers. Step-level interactions are abstracted into canonical events, then organized into a unified graph with three edge types: semantic compatibility, transition dynamics, and outcome evidence. A contrastive spectral decomposition isolates stable, success-aligned behavioral modes while suppressing failure-prone shortcuts. These modes populate a skill tower with three levels: action templates, procedural routines, and overarching task strategies, continuously refined by verifier-guided feedback.
The results are believable precisely because they're not huge. 87.31% success on ALFWorld at 10.35 steps means the agent finishes nearly 9 of 10 household tasks without wandering. The 0.26 invalid actions per task is the quiet number: fewer than one wrong action per task on average, which over a long session saves a lot of context budget. 50.67% exact match on WebShop, where the agent must pick the exact correct item from a long product list, is a strict bar. Half of sessions ending with the right pick is respectable.
DeerFlow: an agent harness with an ops story
ByteDance's DeerFlow hit #1 on GitHub Trending in February 2026, right after the 2.0 launch. The star count isn't the interesting part. Two details matter: 2.0 is a ground-up rewrite with no shared code with v1, and the maintainers rebuilt the project as a product. Sub-agents, memory, and sandboxes sit under one harness, driven by extensible skills. A desktop tool called LLM Space lets you prototype agent ideas, inspect each harness step, replay failures, and benchmark performance.
The setup flow tells you who this is for. You can hand a coding agent a one-line instruction to clone the repo and follow Install.md, which is a neat way to run an agent to build an agent. The make setup wizard walks through LLM provider, search, sandbox mode, and file-write tools, then generates a minimal config.yaml in about two minutes. make doctor verifies the setup. make support-bundle writes a redacted issue summary and a draft issue file for AI-assisted bug reports, excluding .env, raw messages, and user file contents.
I ran into the Docker Compose wall on my first try. Anything older than Compose 2.24 fails to parse the optional env_file syntax in the dev compose file, and the error message doesn't tell you why. Two minutes with make doctor fixed it. That kind of operational polish, a support bundle that writes a usable issue draft, is what separates a harness people actually use from a research demo.
Deployment sizing translates directly to your planning:
| Deployment target | Starting point | Recommended | Notes |
|---|---|---|---|
| Local evaluation / make dev | 4 vCPU, 8 GB RAM, 20 GB free SSD | 8 vCPU, 16 GB RAM | Fine for one developer with hosted model APIs. 2 vCPU / 4 GB is usually not enough |
| Docker development / make docker-start | 4 vCPU, 8 GB RAM, 25 GB free SSD | 8 vCPU, 16 GB RAM | Image builds and sandbox containers need more headroom |
| Long-running server / make up | 8 vCPU, 16 GB RAM, 40 GB free SSD | 16 vCPU, 32 GB RAM | For multi-agent runs, report generation, heavier sandbox workloads |
If you plan to host a local LLM on the same box, size that separately. And multi-worker production requires Postgres plus the Redis stream bridge. Process-local event stores can't enforce singleton delivery receipts across workers.
Quick take: The pattern across the week's releases is uniform: agent quality is becoming an architecture problem, not a prompting problem. Whoever structures the agent loop best wins.
First Things First: make must-haves actually mandatory
The FTF-rl paper starts from a simple question: can agents handle requirements that conflict? The authors construct 3,649 problems across e-commerce, booking, map-based, and ride-hailing scenarios, under three regimes:
- Must-haves uniquely determine the answer.
- Multiple answers satisfy must-haves, and nice-to-haves break the tie.
- No answer satisfies must-haves, so the agent should abstain.
Current MLLMs fail catastrophically in all three. They misinterpret requirements, violate must-haves to satisfy nice-to-haves, and produce invalid solutions instead of abstaining. The fix, First Things First RL, explicitly optimizes reasoning over multi-priority requirements. It also generalizes to LogicVista, MathVision, and InfoQA, which suggests the skill is broadly applicable, not a benchmark hack.
The abstention result is the one to internalize. In scenario three, a model that says "I can't satisfy your hard constraints" is more useful than one that hallucinates a close-enough answer. Most agent scaffolding gives the model no escape hatch, and the model invents one, usually badly.
DMoA: structured debate for clinical calls
Single-turn question answering doesn't reflect how diagnosis works. Clinicians don't ask one question and accept one answer. They argue, weigh differentials, revisit findings. The Debate-Mixture-of-Agents framework maps that process onto role-based interaction: agents take positions, iterate, and converge, evaluated across two datasets totaling 2,016 cases (297 rare diseases, 1,719 challenging cases).
The gains over GPT-4o: 10.21 percentage points on most-likely diagnosis accuracy, 11.36 on safety. The ablation matters more. The improvement wasn't from throwing more models or longer outputs at the problem. A 4x2 structure (four agents, two rounds), stronger base models, and bigger token budgets all helped, but the structured workflow itself drove a substantial share of the gain. That's the part to steal for non-clinical work: role separation and iterative rounds beat prompt ensembling.
AutoHedge: agents with a risk gate
The Swarm Corporation's AutoHedge is an autonomous agent hedge fund, currently trading Solana, with Coinbase on the roadmap. The architecture is a linear pipeline, and the order of stages is the design: the Director Agent generates strategy and theses, the Quant Agent does technical and statistical analysis, the Risk Management Agent handles position sizing, and only then does the Execution Agent act.
The order matters more than any individual agent. Position sizing gates execution. That's the difference between an interesting experiment and a system you'd trust with money, and it's the piece most hobbyist trading bots skip. Structured JSON outputs and detailed logging are table stakes. The risk-first ordering is the actual thesis.
Getting it running was simple enough: pip install, then a Jupiter API key for price data, plus a wallet private key for trading. The scope check matters more than the setup. Solana only for now, Coinbase later. The README frames it as institutional-grade, but the community conversation leans more grounded: experiment infrastructure first, returns second.
The pattern: architecture is the agent
Across five systems, the same structural fixes appear at different layers:
- Trace2Tower structures memory: skills, not raw trace dumps.
- FTF-rl structures reasoning: priorities, and an explicit abstain path.
- DMoA structures interaction: roles and iterative rounds.
- DeerFlow structures the runtime: sandboxes, memory, and an ops flow that makes failures reproducible.
- AutoHedge structures risk: a mandatory gate between analysis and execution.
None of these are about a better model. They're about constraining where the model's output can go and what it can build on. For anyone who needs reliable results, the weeks of "just prompt it harder" are over.
Common pitfalls
The mistakes here are the ones I keep seeing in agent codebases, including my own.
Treating traces as flat context
Copy-pasting raw trajectories into the prompt bloats the context and loses temporal dependencies. Trace2Tower exists because flat retrieval fails. If you're doing skill induction, cluster by outcome-conditioned structure, not lexical similarity. Your retrieval index should know which paths led to success and which led to loops.
Letting nice-to-haves drown out must-haves
FTF-rl's evaluation shows catastrophic failures the moment requirements conflict. Separate hard constraints from soft preferences explicitly, and give the model an abstain route when nothing satisfies the hard constraints. A refusal is a valid answer. A hallucinated near-miss is not.
Expecting single-turn Q&A to carry diagnostic workloads
DMoA's ablation shows the gains aren't from more models or longer outputs. If you add agents without roles and rounds, you pay token costs and get ensemble averaging. Assign positions, force iteration, measure safety and accuracy separately.
Running the DeerFlow stack on an old Docker Compose
Compose older than 2.24 fails to parse the dev compose file, and the error doesn't tell you why. Check the version before opening an issue. Also: 2 vCPU / 4 GB will not run even a light session. And if you scale past one Gateway worker, you need Postgres and the Redis stream bridge, or delivery receipts will silently break across workers.
Executing before risk checks
AutoHedge's ordering exists for a reason. In any agent that spends money, writes files, or calls external APIs, put a risk or approval gate between analysis and execution. Position sizing before order placement. Permission checks before file writes. The agent can be brilliant and still cause damage.
One thing to remember: All five systems treat agent quality as an architecture problem. The model still matters, but this week's wins came from how the loop is structured: which memories persist, which requirements outrank others, which agents get to speak, and what gates execution.
The bottom line
Three scenarios to take home:
- If you're building agents that learn from experience, adopt hierarchical skill induction like Trace2Tower instead of flat trajectory retrieval. Outcome-conditioned skills give you reusable routines at a fraction of the context cost, roughly 10 steps per ALFWorld task with under one invalid action.
- If you're building agents for constraint-heavy or high-stakes domains (clinical, booking, finance), add explicit priority signals and structured debate. The DMoA ablation is the evidence: 10.21 points of accuracy came from workflow structure, not model count.
- If you're deploying a production harness, copy DeerFlow's ops playbook: setup wizard, doctor checks, support bundles, sandbox modes. Budget at least 4 vCPU / 8 GB for a single light session, and plan for Postgres plus Redis before you touch multi-worker.