Appearance
The demo works, the deployment doesn't
Every agent demo looks good in the first session. The model solves the task, calls the right tools, writes the answer. Stretch the same task over days, across replans and handoffs, and it falls apart. Objectives drift, evidence goes stale, context balloons, and nobody can tell what the agent actually did.
A cluster of recent papers and projects makes the cause clear. It isn't the model. The harness around the model decides whether long-horizon work survives contact with reality.
The harness is where the accuracy lives
Start with Bioinfoysis, a multi-agent harness for bioinformatics. The authors ran the same four LLMs through their harness and measured the accuracy lift. On SeqQA2, average accuracy went from 27.81% to 64.13%. On DbQA2, from 3.13% to 31.25%. Same models. Different scaffolding. On BixBench, the harness hit 82.4%, the best published result on that benchmark.
Think about what a 2.3x lift means. Upgrading a model generation rarely buys you that much. A harness that governs planning, execution, memory, and evidence flow did, across four different base models.
The CAE simulation paper is even more blunt. On FoamBench, a plain single-agent harness running on a modern model scored 96.4%. A specialized multi-agent CAE system scored 88.2%. The paper asks what a CAE agent needs beyond a generic harness. Its answer: almost nothing from the agent playbook, and everything from domain knowledge.
| Configuration | FoamBench accuracy |
|---|---|
| No repair round | 71.8% |
| Scripted reflection | adds nothing |
| Specialized multi-agent CAE system | 88.2% |
| Single-agent harness with execution-feedback repair | 96.4% |
| Single-agent harness with solver tutorials | 96.4% |
Execution-feedback repair (let the agent see the error and fix the script) is worth 24.6 points. Scripted reflection is worth zero. Domain knowledge delivered as solver tutorials is the single largest gain measured in the whole study, lifting accuracy from 80.9% to 96.4%.
That inverts two years of agent orthodoxy. Multi-agent decomposition, domain retrieval, and reflection were invented to compensate for weak base models. Modern harnesses already give you multi-turn reasoning, tool use, and execution feedback. The remaining levers are evidence discipline and domain knowledge.
Evidence flow is a silent killer
Bioinfoysis's design explains why harnesses matter. Most agent systems treat planning, tool use, and code execution as transient interactions that exist only to produce a final answer. Bioinfoysis represents each request as a persistent, artifact-grounded analysis run. The planner keeps an executable checklist and revises pending steps using structured handoffs returned after each worker execution. Those handoffs bind every intermediate result to the agent, checklist step, and plan generation that produced it.
The binding is the point. When the plan changes, evidence produced under the old plan can't be silently reused. A controlled runtime validates generated scripts, tables, and figures before anything downstream consumes them.
This is the failure mode I keep hitting in my own research agents. The model replans, and two turns later it cites a figure produced before the replan. The conclusion looks grounded. It isn't. Artifact-based validation is the fix, and it's boring enough that most agent frameworks skip it.
Quick Take: Same models, different harness, a 2.3x accuracy jump. For long-horizon work, the bottleneck is scaffolding, not the LLM.
Skills are the new packaging unit
If the harness is the runtime, skills are how expertise ships. The last few weeks brought a flood of installable skill libraries. Google released google/skills, with hundreds of agent skills for Google Cloud: GKE configuration, BigQuery, monitoring queries, cloud architecture. text-to-cad packages the entire CAD/CAE/CAM workflow as skills, from STEP model generation to g-code slicing to SendCutSend file checks. MathModelAgent turned a math-competition agent into a pure skill layer with 17 Typst paper templates and a 9-step verification pipeline.
| Project | Domain | What it packages |
|---|---|---|
| google/skills | Cloud operations | Official skills and plugins for Claude Code, Codex, Antigravity CLI |
| text-to-cad | CAD, CAE, CAM | CAD generation, DXF, URDF/MoveIt, g-code slicing, printability checks |
| MathModelAgent | Mathematical modeling | End-to-end pipeline, 17 Typst templates, evaluator feedback loops |
| SimSkill | Traffic simulation | Self-evolving skill library over SUMO, up to +25 points verified completion |
SimSkill is the research version of the same idea. It identifies capability gaps, generates and solves environment-grounded tasks, verifies solutions through an action-critic loop, and consolidates what works into episodic, procedural, and semantic memory. No backbone model updates. On held-out benchmarks it improves verified completion by up to 25 percentage points. The honest caveat: memory doesn't help every model, and it doesn't uniformly cut inference cost. The value is backbone-dependent.
The MathModelAgent README tells a story I keep hearing. Two years ago, the author maintained a full multi-agent framework. Now the project is skills-driven: you install it into Claude Code or Codex and run one command. The author's point, in his own words, is that building on harnesses plus skills beats maintaining a bespoke framework. I've reached the same conclusion in my own projects. The framework is a commodity now. The expertise is not.
Control planes keep state out of the chat
Skills solve knowledge. State is a separate problem, and loopx is the clearest articulation of it so far. loopx is an open, provider-neutral control plane that runs on top of any agent harness: Codex, Claude Code, Cursor, or your own runner. Objectives, gates, todos, evidence, quota, and handoffs stay durable while the harness executes bounded turns.
The mental model is an agent-native Kanban. Cards carry identity, authority, evidence, and continuation. Moves are validated operators: claim, gate, monitor, writeback. The dashboard is a projection. The loopx state is the source of truth.
The notable thing about loopx is what it refuses to be. It is not an autonomous production controller. Publishing, production writes, dangerous permissions, and final ownership stay with a human. The public showcases span 200+ hours of wall-clock loop lifetime, with one user reporting four days unattended and another attributing seven merged PRs to a loopx-governed refactor. Those claims are user-reported and hedged accordingly. The direction is still clear: the control state outlives any single session, restart, or even a switch between harnesses.
The next shift: agents that grow their own structure
Two further signals point past current harness design. The proactive service agents survey formalizes what most agent builders do informally: an agent must decide whether to stay silent, ask, assist, or act, while pricing interruption, misunderstanding, overreach, and privacy costs. Its sharpest conclusion is that offline classification performance doesn't predict deployment benefit, and long-term memory is not what makes an agent proactive. What matters is calibrated intervention value, verifiable authorization, recoverable execution, and counterfactual evidence.
Eureka, a meta-architecture from a university collaboration, goes further. Instead of a fixed agent structure, it dynamically compiles the task into an obligation graph, spots architecture hotspots, and grows dedicated macro-agents with their own memory, tools, and verification standards. In testing it finished all 170 recursive long-horizon tasks, produced 3,948 verifiable certificates, and showed zero false task completions. The headline: on a known proof route toward the Riemann hypothesis, it extended the validity of a key certificate from a ≤ 1/4 to a ≤ 69/200 (0.345), covering 99.55% of the first theoretical threshold. Short of a proof, but a verifiable milepost on that path.
The engineering numbers matter just as much. Separating semantic planning from deterministic computation cut the core workspace's context load by 57.8% and avoided 65.38% of repeated computation. In practice that means the same token budget holds more evidence and each step costs less. A 16,000-task concurrency test produced results identical to sequential execution, which is the precondition for trusting parallel agents at all.
Key numbers across this cluster
| Metric | Value |
|---|---|
| BixBench accuracy, Bioinfoysis | 82.4%, best published |
| SeqQA2 accuracy lift across 4 LLMs | 27.81% → 64.13% |
| DbQA2 accuracy lift across 4 LLMs | 3.13% → 31.25% |
| FoamBench gain from execution-feedback repair | 71.8% → 96.4% |
| FoamBench gain from solver tutorials | 80.9% → 96.4% |
| SimSkill verified completion gain | up to +25 points |
| Eureka context load reduction | 57.8% |
| Eureka repeated computation avoided | 65.38% |
Common pitfalls
Five things trip people up repeatedly in this space, several confirmed by the CAE ablations.
Don't reach for multi-agent before proving single-agent fails. The FoamBench numbers are unambiguous: specialized multi-agent machinery scored 8 points below a single-agent harness. Start with one loop and execution feedback. Add structure when you have evidence the single loop is the bottleneck.
- Don't treat the chat context as memory. A context window is not durable state. loopx exists because objectives, evidence, and handoffs must survive restarts, harness changes, and days of wall-clock time. Write state down, version it, keep it outside the conversation.
- Don't add scripted reflection and expect gains. The CAE study measured it: reflection added nothing beyond what execution feedback already provides. The model already sees its errors when failures are fed back. A prompt asking it to be more careful is not an improvement.
- Don't let replanning invalidate evidence silently. Bind each artifact to the plan generation that created it, the way Bioinfoysis does. Silent reuse of stale tables and figures is the most common way a long agent run produces a confident wrong answer.
- Don't give the agent irreversible actions without a policy gate. The strongest community critique of Needflare pushed exactly this: let the model produce structured decisions, but let a policy engine validate severity, resource constraints, and authorization before anything irreversible happens. Add idempotent event processing and a confidence threshold that downgrades to rule-based triage when the model is unsure.
What the community is saying
The discussion around Needflare is worth reading even if you never touch disaster response. The most useful critique I found argued that the real innovation isn't combining three Google models, but designing graceful degradation around unreliable infrastructure: local-first ingestion, on-device PII scrubbing, and async reconnection. The same commenter wanted replayable event logs, regional failover, and explicit model confidence thresholds. Rules before autonomy, in other words. I've hit the same wall in my own agent pipelines: the model makes a structured decision, and nothing between it and the irreversible action validates the severity, the authorization, or the duplicate.
On the skill libraries, the recurring question is versioning, and I found it on the first install. npx skills add overwrites whatever is already installed, while npx skills update only refreshes skills already in your lockfile, so it silently misses new ones. If skills are becoming the package manager for agent expertise, we're going to need the discipline of the old one: lockfiles, semantic versions, and retirement paths.
The appetite for learning is equally loud. The Datawhale hello-agents course, a free 16-chapter path from building ReAct by hand to training agents with Agentic RL, passed 70,000 stars. Its framing splits the field into flow-driven platforms like Dify and Coze versus AI-native agents, and the community has voted decisively for building from first principles.
One thing to remember
Every domain in this cluster shipped the same lesson in different packaging. Bioinformatics needed artifact-grounded runs. CAE needed execution feedback and tutorials, not multi-agent theatrics. Traffic simulation needed a reusable skill library. Long-running engineering work needed durable control state. The model is a necessary ingredient, and it is no longer the interesting one. The harness, the evidence graph, and the skills library decide whether your agent finishes the job on day three or collapses on turn twelve.
The Bottom Line
If you're building an agent that runs for hours or days across tools, adopt artifact-grounded runs with validated intermediate outputs and structured handoffs. Stale evidence is the silent killer; bind every result to the plan that produced it.
If you're choosing between a bespoke multi-agent framework and a modern harness plus installable skills, take the harness. Execution feedback and domain tutorials outperformed multi-agent machinery in the only head-to-head numbers we have, and skill libraries are absorbing domain expertise fast.
If a human carries final responsibility, put a control plane between the agent and anything irreversible. Explicit gates, quota, and typed handoffs keep long-running work reviewable and restartable, and the agent should never hold the authority for publishing, spending, or final ownership.
The throughline across all of this is that long-horizon autonomy is a systems property. You don't prompt your way to it. You design for it: durable state that outlives sessions, evidence that stays bound to the plan that created it, and skills that accumulate value across every run. That's the difference between an agent that demos well and one you'd trust with a week of your work.