Appearance
Agents got real. The tooling is catching up.
Google Duplex booked a hair appointment over the phone in 2018. It was a constrained RNN, trained on a closed domain, and it felt like a magic trick. The system hit under 100ms response latency on simple turns, fast enough that the person on the other end never noticed they were talking to software. Eight years later, the same company positions Gemini 3.7 Flash explicitly for "complex agentic tasks at scale," orchestrating sub-agents the way Duplex orchestrated a single phone call. Xero has agents managing multi-week tax workflows. Shopify runs parallel subagents over long-horizon merchant data. Databricks uses agentic loops to diagnose pipeline failures. The magic trick became a product category.
The hard problems shifted with it. Nobody's asking whether frontier models can chain tool calls anymore. The questions are operational: how much memory should an agent carry, how do you see what it's doing while it runs, and how do you trust the work when it's done? Three recent developments point at the same answer. IBM's memory calibration study, a trace dataset mined from Claude Fable-5 agent runs, and the AdaLens oversight system all converge on one conclusion: the agent infrastructure layer is taking shape, and it has three pillars. Calibrated memory, trace data, and interactive observability.
Memory is a dose, not a feature
IBM Research's ALTK-Evolve work starts from a simple idea: distill lessons from an agent's past trajectories into guidelines, inject them back at inference time, and the agent gets better. No weight updates, no human annotation. When they scaled this to eight models on AppWorld, 585 multi-step tasks across 9 simulated apps, the simple idea broke.
The finding: agentic memory is a dose you calibrate to the model. Three patterns showed up across the capability spectrum.
| Model | Pattern | Baseline TGC / SGC | Best-memory TGC / SGC | Best config | Δ TGC | Δ SGC |
|---|---|---|---|---|---|---|
| gpt-oss-120b (117B MoE) | Weak / selective | 39.9 / 21.4 | 56.0 / 37.5 | curated retrieval | +16.1 | +16.1 |
| DeepSeek-V3.2 (671B MoE) | Strong w/ headroom | 79.8 / 64.3 | 89.3 / 80.4 | full guideline set | +9.5 | +16.1 |
| Claude Opus 4.6 | Strong w/ headroom | 90.5 / 87.5 | 94.6 / 94.6 | full guideline set | +4.1 | +7.1 |
| GPT-5.5 | Strong (near-ceiling) | 92.3 / 82.1 | 95.2 / 89.3 | full guideline set | +2.9 | +7.2 |
| GLM-5 (745B MoE) | Saturated | 87.5 / 80.4 | 87.5 / 80.4 | full guideline set | 0.0 | 0.0 |
Strong models with headroom want the full guideline set, including rare edge-case lessons. They have the capacity to absorb all of it. DeepSeek-V3.2 climbed +9.5pp on task completion when given everything. Weaker models drown in a large set; gpt-oss-120b gained +16.1pp from a compact core plus a few task-relevant guidelines, while the full set gained less and cost more. And already-saturated models show no measurable gain at all. GLM-5 sat flat at 87.5 / 80.4 no matter what was injected.
The stricter SGC column matters more than it looks. SGC is all-or-nothing per scenario: the agent has to clear every variant of a task, not just the average case. Memory gains are consistently larger there, which means guidelines improve reliability, not just mean performance. DeepSeek's SGC jumped +16.1pp against a +9.5pp TGC gain. Even GPT-5.5 and Claude Opus, both near the ceiling on TGC, gained +7.2 and +7.1pp SGC respectively. Memory keeps paying off as long as a model has a remaining failure mode to target.
The cheapest memory strategy is often the best
The obvious objection to injecting guidelines is cost. The full set gets re-sent on every ReAct step, and the numbers show why that matters.
| Model | Config | Tokens/task (baseline) | Tokens/task (+ memory) | Overhead |
|---|---|---|---|---|
| DeepSeek-V3.2 | full guideline set | 148K | 263K | +78% |
| gpt-oss-120b | full guideline set | 110K | 166K | +51% |
| gpt-oss-120b | curated retrieval | 110K | 116K | +5% |
Curated retrieval wins on accuracy and cost at the same time: +16.1pp TGC at only +5% tokens. Better performance without a bigger inference bill. And memory doesn't lengthen trajectories. DeepSeek runs about the same number of ReAct steps with or without memory, around 18 to 19 on average. The added cost is input-token inflation, not longer reasoning loops.
The production lever is prompt caching. The static portion of a guideline set is identical across steps, so keep that shared prefix stable and it stays cacheable. Cache-aware prompt design turns a +78% token overhead into a rounding error.
Quick Take: Agent memory is a hyperparameter, not a feature. The right dose depends on the model's headroom, and for weaker models the selective option is both more accurate and cheaper.
Observability with a steering wheel
The AdaLens paper starts from a specific pain: long-running agentic data analysis. An agent runs for hours, spawns parallel branches, tries a hypothesis, abandons it, deepens another. Conventional chat interfaces show you a transcript, not a decision structure. You can't tell what the agent is doing, and you can't redirect it before it burns compute on a dead end.
AdaLens builds a storyline: a unified view that ties together the analytical plan, execution progress, intermediate findings, and which data columns the agent has touched. Steering is grounded in those elements. You can push the agent toward a promising branch or pull it out of a low-value one, mid-run.
The design constraint is the part worth stealing. Existing interactive tools were built for discrete, turn-by-turn exchanges. Long-running agents don't work that way. They branch, backtrack, and revise their own plans. AdaLens treats oversight as an ongoing interaction with a live system, not a post-mortem on a log.
Agent traces are becoming training fuel
A dataset landed on Hugging Face in late July: 12,730 records of Claude Fable-5 agent traces, cleaned and released under MIT. It's an SFT dataset with a stated priority of quality over quantity, and it ships in two formats. The standard OpenAI chat format works with Axolotl, Unsloth, and the OpenAI fine-tuning API. The agent_traces format preserves the user/assistant/tool role structure, plus a separate reasoning field for models that support explicit thinking tokens.
| Property | Value |
|---|---|
| Total records | 12,730 |
| Train split | 5,728 (45.0%) |
| Validation split | 318 (2.5%) |
| Test split | 319 (2.5%) |
| License | MIT |
The size isn't the point. 12,730 records is enough for a solid SFT run on a single node. Fine-tune a 3B model and you have a coding agent. The bigger shift: agent traces are becoming a commodity input, the same way instruction data did in 2023. The quality distribution tells a different story. 6,740 of 12,730 records sit in the 0.9 to 1.0 quality band, and almost nothing sits below 0.5. Someone already learned the lesson that raw traces teach agents to be sloppy.
When agents audit papers, humans still steer
The ICML 2026 Open Reproductions challenge is the largest attempted reproduction of a scientific conference: 1,221 participants, 2,962 cloud jobs, and a conference that accepted 6,352 papers, roughly double the previous year. The headline numbers are stark.
Key Numbers
- 51% of examined papers had at least one claim independently verified
- 23% had at least one claim falsified or contested
- 242 papers where independent teams reached opposite verdicts on the same claims
- 2,962 cloud jobs launched by 1,221 participants
Some findings are brutal. A spotlight paper on learning-augmented paging claimed a robustness bound that a participant's logbook broke at a specific step. The additive term grows with k, and re-implementation at k = 1,024 confirmed the failure at roughly nine sigma. The reviewer's own note admitted, "I did not check all the proofs carefully." Three independent teams found counterexamples to a token-collapse theorem, with violations first appearing at t = 224, ~3,800, and 6,416 steps. That neatly explains why everyone else verified it: finite-horizon checks stop too early. Another paper's entire theory section analyzed reverse KL divergence while the released code computed forward KL. And a transformer cache paper's headline "3.1% quality cost" became roughly 9.4% once a participant noticed that about 66% of evaluated label positions were EOS padding tokens.
When I ran my own reproduction attempt, the pattern was consistent. Agents got stuck in local loops, misread scale-dependent behavior, and occasionally built an entire falsification on top of a units mismatch. One logbook claimed a method was 2x slower than baseline; it turned out to be a per-trajectory versus per-batch-of-50 arithmetic error, and the participant's own data confirmed the paper's claimed 8x speedup. The reliable results came from workflows where a human was steering: re-pointing the agent, questioning an assumption, deciding an experiment's premise was wrong before burning a week of compute on it. I've also been following a complementary project that runs clean-room reproductions under roar, which passively captures provenance instead of relying on code instrumentation, so every attempt produces an AI-BOM and a reproduce path anyone can re-run. The tooling is getting better. The human is still the bottleneck.
Common pitfalls
Dosing memory without checking model headroom. Dumping the full guideline set into a weaker model doesn't just fail to help, it costs. gpt-oss-120b gained less from the full set than from curated retrieval and paid ~50% more tokens for the privilege. Calibrate before you inject.
Trusting verification that stops too early. The paging paper collected multiple "verified" verdicts because checks stopped before the growth became visible. When you audit agent work, ask what horizon the check actually covered. t = 224 matters.
Treating traces as free data. Raw agent traces teach agents to be sloppy. The fable-5 dataset is rigorously cleaned for a reason. If you're mining your own traces, filter hard, and keep the reasoning field separate from the assistant content so models with thinking tokens can use it.
Ignoring token inflation in memory injection. A full guideline set re-sent every ReAct step costs +78% tokens on DeepSeek. Design for prompt caching from the start: keep the static guideline prefix stable so it stays cacheable.
No steering until post-mortem. Long-running agents branch and backtrack. If your only observability is a transcript after the fact, you've already spent the compute. You need intervention points mid-run.
One thing to remember
The through-line across all of this, memory calibration, trace datasets, storyline observability, adversarial reproduction, is that the model stopped being the interesting variable. Every gain in this article came from changing what the agent sees, how it's monitored, and how its work is checked. Not from a bigger model.
The Bottom Line
If you're building long-running agents, adopt storyline-style observability with mid-run steering. AdaLens shows the pattern: unify plan, progress, and findings into one view, and give the operator directional controls. Post-hoc logs are too late.
If you're adding memory to an agent, calibrate the dose to the model tier. Weaker models want a compact core plus per-task retrieval, which is both more accurate and cheaper. Frontier models with headroom want the full guideline set, kept affordable via prompt caching. Saturated models want nothing until their failure modes are understood.
If you're fine-tuning on agent traces, filter like your model's behavior depends on it, because it does. And expect two things within the next year: trace datasets will standardize the way instruction datasets did, and agent-based verification will become a standard part of the research pipeline, with humans still doing the steering.