Appearance
Three trust boundaries for safety-critical LLM agents
The trust problem
LLM agents stopped being read-only. Claude Code, Codex CLI, and Gemini CLI all run local tools by default. Security teams are testing agents as autonomous SOC analysts. And the agent's own tool path will happily overwrite a file, run bash, or push a containment action, with the bytes landing before anyone reads the audit trail.
The trust model here is inverted from traditional security. You don't vet the binary, you vet a runtime that writes files based on a prompt. The skill you installed might carry hidden instructions. The model might recommend a containment action that contradicts the network topology it's reasoning over. The hook that was supposed to block a write might only log it.
Three trust boundaries have crystallized in the last few months, and each one has a recent project that shows how it's done:
| Trust boundary | When it fires | Example tool | Failure mode it closes |
|---|---|---|---|
| Install time | Before a skill is loaded into the agent | NVIDIA SkillSpector | Skills carrying hidden instructions or malicious patterns |
| Decision time | When the agent selects an action | Sentinel-RL | An action that contradicts the graph topology it runs on |
| Effect time | Right before a write lands on disk | gx escrow hook on OpenClaw | Bytes landing before anyone can refuse them |
The three projects don't compete. They layer. And taken together, they define what "safe agent" is starting to mean.
Sentinel-RL: the model narrates, the policy decides
The Sentinel-RL paper takes on a specific failure: an LLM agent acting as a SOC analyst at enterprise scale. Two limitations make that unreliable. A context window cannot hold a multi-thousand-host authentication graph, and free-form generation gives no guarantee that a recommended containment action is consistent with the topology it operates on.
The fix is architectural, not prompt-based. A heterogeneous graph attention encoder summarizes the live authentication subgraph into a fixed-dimensional state. A Proximal Policy Optimization agent maps that state to a constrained set of investigative actions. The LLM is restricted to consuming the policy's recommendations and producing analyst-readable narratives, gated by a critic.
The key move is that the model never composes the action. It narrates choices the policy made. If the model hallucinates, the worst case is a bad paragraph, not a bad containment decision.
The numbers hold up at enterprise scale. The system loads a 24M-edge authentication subgraph into Neo4j in 14.2 minutes on a single 32-core node, about 24x faster than the canonical MERGE-based pipeline. 24M edges is roughly what a mid-size enterprise generates; 14.2 minutes means you can refresh the graph between shifts instead of weekly. The sliding-window alert engine trips a 25-event/10-second threshold in under 2.5 seconds across 50 trials, which keeps a 10-second detection SLA honest.
Key numbers: 24M-edge auth graph loaded in 14.2 min, 24x faster than MERGE. Alert engine trips a 25-event/10-sec threshold in ≤2.5 s. PPO converges to mean episodic return 8.74 ± 0.31. Held-out precision 0.91, recall 0.87 on red-team events. Full detect-investigate-recommend-approve loop: 6.3 s median.
PPO converges to a mean episodic return of 8.74 ± 0.31, and the tight variance is the practical signal: the policy is stable, not still drifting while it's on duty. On labeled red-team events, precision of 0.91 and recall of 0.87 mean it misses roughly one attack in eight and burns an analyst's time on false alarms about one time in eleven. The integrated loop completes the full detect-investigate-recommend-human-approve cycle at a median of 6.3 seconds, and that number makes the human-approval boundary believable. Approval stays in the loop without stalling it.
This is the decision-time boundary, solved by removing the decision from the language model.
The one place a write can still be stopped
The OpenClaw write-path article is the effect-time story, and it starts from a blunt observation: an agent using OpenClaw can decide to overwrite a file, and by the time anyone finds out, the bytes are already on disk. A log entry written after the fact doesn't help. The author wanted one point in the tool path where a write could still be refused before it lands.
OpenClaw has exactly one such point: the before_tool_call hook. Everything before it can still say no; everything after it is already history. Two documented facts make the seam usable. runBeforeToolCall runs sequentially, can block, and can rewrite the call's params before the tool sees them. runAfterToolCall is fire-and-forget. By the time it runs, there's nothing left to hold onto.
So the author built a plugin that sits at that seam and puts every proposed filesystem write through gx, an escrow-and-inverse layer. The flow has four verdicts:
The Unknown branch is the one worth quoting. Blocking with "couldn't reach the decision service" is not the same as blocking with "policy denies this." Folding "couldn't ask" into "asked and no" is the one shortcut a reversibility layer can't take without lying about the thing it exists to be honest about. This matters because the block reason lands in the model's context, and an agent handed a fake policy denial learns that the guardrail is an oracle instead of a service with outages.
To preempt the obvious objection, the author ran the plugin inside a real OpenClaw gateway with a DeepSeek-driven agent turn. The Admit, Deny, and Escalate verdicts were exercised against a real write, a real attempt to overwrite /etc/hostname, and a file exceeding the inverse-construction limit. Each verdict was checked against the signed gx receipt independently, and an admitted write was restored byte-for-byte with gx undo. It's also honestly scoped: 14 stars, 4 forks, a v0.1.0-alpha tag, and only the write tool wired at the time. The seam demonstrably works. It doesn't follow that it always will.
Quick Take: a write path is only as safe as the last seam that can still say no after the model has finished thinking.
The review that changed the design
The most useful part of the OpenClaw post turned out to be the comment thread. A reviewer walked the Admit branch and found a hole that changed the design. The exchange shows how a write-path safety layer actually fails.
The sharpest objection went straight at the Admit branch. gx does the disciplined part: submit, plan, verify, commit, with the plan snapshotting a precondition fingerprint and the commit rechecking it before applying. But the handler returns { params }, and OpenClaw's native write tool runs after the hook has already yielded. When I traced that path end to end, the gap was concrete: the receipt proves the gx transform was valid, but nothing proves the bytes now on disk descend from the receipted post-image. Another process, a later hook, or the agent's own unmediated bash could change the file between the gx commit and the native write.
The deeper issue was hook ordering. runBeforeToolCall is sequential and can rewrite params, so a hook registered later can rewrite the params after gx has already escrowed the old content. If gx escrows content A, produces post-image B, and a later hook rewrites the params to C, the receipt and the inverse both describe a write that never reached disk. Undo would restore a pre-image for a transition that did not happen. That's worse than a gap on the forward path, because the recovery path lies to you.
The author checked whether "am I last for this tool" is even answerable. It isn't. The hook result type has no field for replacing the result, and params is merged lastDefined across the chain, so a later plugin overwrites yours by design. Rewritten params are a request, not a guarantee. The fix took the stronger form: Admit now returns block: true. gx applies the write itself and refuses the native call. No bytes move without going through the membrane, and there is no second write left to be last in front of. The test suite went from 7 failures to 0, and reverting the patch brought all 7 back.
The bug is less worrying than the pattern behind it. The author found seven places in his own gates where a comment asserted that a range was checked when the range wasn't. Comments that say "not an omission" read as though they were verified. Nothing re-derives them. If you build safety layers for agents, go find your own seven before someone else does.
SkillSpector: scanning the inputs before they run
The decision-time and effect-time boundaries assume the agent itself is what you trust. But agents today ship with skills: SKILL.md files and companion scripts that Claude Code, Codex CLI, and Gemini CLI execute with implicit trust and minimal vetting. NVIDIA's research puts the problem in numbers: 26.1% of skills contain vulnerabilities and 5.2% show likely malicious intent. One in four installed skills has something wrong. One in twenty is hostile. That's the install-time boundary.
SkillSpector is a scanner that answers a single question: is this skill safe to install? It scans git repos, URLs, zips, directories, or a single SKILL.md, and runs 71 vulnerability patterns across 17 categories: prompt injection, data exfiltration, privilege escalation, supply chain, excessive agency, output handling, system prompt leakage, memory poisoning, tool misuse, rogue agent, anti-refusal, trigger abuse, dangerous code via AST, taint tracking, YARA signatures, MCP least privilege, and MCP tool poisoning.
| Pattern family | Example patterns | Severity |
|---|---|---|
| Prompt injection | P1 instruction override, P5 harmful content, P9 whitespace padding | MEDIUM to CRITICAL |
| Anti-refusal | AR1 refusal suppression, AR3 safety policy nullification | HIGH |
| Exfiltration | E2 env variable harvesting, E4 context leakage | MEDIUM to HIGH |
| Privilege escalation | PE2 sudo/root execution, PE3 credential access | LOW to HIGH |
Two architectural choices stand out. First, two-stage analysis: fast static analysis always, optional LLM semantic evaluation when you want it. The scanner reports scan_mode and llm_used so a low score from static-only analysis is never mistaken for a clean full scan. Second, the ingest caps fail closed: 100 MiB per remote input and 10,000 zip members, with a breach raising IngestLimitExceededError. A zip bomb gets rejected before it lands on disk, not after.
The practical pain points show up in the details. The stdio MCP transport still has a known initialize hang (issue #199), so the first scan can stall if your client isn't patient. The batch scanner was originally built against OpenAI-compatible endpoints, and DeepSeek's lack of structured-output support required manual JSON-parsing patches. None of this stops the core flow. All of it tells you this is a fast-moving tool with rough edges.
The MCP server is the interesting part. skillspector mcp exposes scan_skill as a tool, so an agent can gate its own skill installs at runtime. That turns SkillSpector from an out-of-band audit into a runtime guardrail, which is exactly the direction the whole cluster is moving. One warning from the README: the HTTP transport ships without authentication, and local paths are rejected over HTTP for that reason. Don't bind it to a routable interface without a proxy in front.
The offensive side is automating too
The defensive tooling exists because offense got cheap. That's the context you need for the exploitarium repo: a consolidated archive of proof-of-concept vulnerability research, much of it produced with an AI-assisted fuzzing workflow. The author is explicit about the division of labor: GPT-5.3 did all the fuzzing, because barely any "thought" is necessary when you have an efficient workflow. The PoCs themselves were hand-typed. The READMEs are openly AI-formatted. And the author is equally explicit that you don't need a SOTA model: the delta is marginal when paired with decent human oversight and a good workflow.
The breadth is the story. 7zip, c-ares, curl, docker, ffmpeg, firefox, gitea, gogs, imagemagick, libarchive, libssh2, nextcloud, nextjs, nmap, qemu, redis, vlc. The long tail of open-source infrastructure gets ground through one automated workflow. The objdump finding alone carries 41 tracked entries; the workflow doesn't just find bugs, it preserves reproductions at scale. And the author's jab at the industry, "a surprising amount of security researchers aren't able to adjust the PoC to work in their environment," is a maturity gap the rest of the field should read as a warning.
OpenAI's $1 billion Daybreak commitment puts frontier cyber AI, training, and support in front of essential services like hospitals, utilities, and transit systems. The scale says the market sees defensive AI as a real deployment problem, not a lab demo. The same capabilities that make automated fuzzing cheap are being pointed at defense, and defense-grade AI has a stricter requirement: it has to be constrainable, which loops back to Sentinel-RL's constrained action space and OpenClaw's write seam. Offense gets to be right once. Defense has to be safe every time.
Common pitfalls
An after-the-fact log entry is not a guardrail; it is a record of damage already done. OpenClaw documents runAfterToolCall as fire-and-forget for a reason. If your write path only instruments the call after it completes, you have forensics, not safety. The seam has to condition the effect, not observe it.
Don't gate installs on a risk score without checking which mode produced it. SkillSpector returns scan_mode and llm_used precisely because a static-only score can look identical in JSON to a full-scan score. If your CI gates on safe_to_install without verifying the scan mode, you've built a threshold that silently means less than it says.
Never fold "couldn't ask" into "asked and no." A guardrail that can't reach its decision service and falls back to denial lies about why it blocked. The block reason goes into the model's context, and an agent that reads a fake policy denial learns that the guardrail is an oracle rather than a service that can fail. When your membrane is down, say so, out loud.
Don't build a safety layer on top of a hook whose ordering you don't control. In a sequential hook chain, a later hook can rewrite your params by design, and lastDefined merging means the rewrite wins. If you can't establish that you're last for a given tool, your layer must not assume final authority over the call. The OpenClaw fix, blocking the native call and applying the write inside the guardrail, removed the second write entirely rather than arguing about ordering.
Finally, treat coverage claims in comments as bugs until proven otherwise. The seven-of-this-shape pattern recurs: something reports a range as checked, and the range isn't checked. A comment asserting coverage is not a test. When you find one, look for the other six.
One thing to remember
Capability is outrunning trust, and the fixes are structural, not conversational. A more careful model does not solve a topology-consistency problem, a skill supply chain, or an unconditioned final write. What solves them is moving the point of refusal away from the model: constrain the action space, gate the write at the seam, scan the skill before it runs. If you can't point to the