Appearance
Public vulnerability data stays on the shelf
Every CVE is a writeup of a mistake that already happened. The databases hold the weakness type, the affected component, and usually a link to the fixing commit: the exact diff that removed the flaw. That's an enormous dataset for training a detector, and almost nobody uses it that way. The records read like documentation for humans, so the knowledge sits static. The same unsafe pattern keeps appearing in unreported code, because the only place it gets caught is where someone already filed an advisory.
BUGSTONE-E2E treats the patch itself as the detector. Instead of a human reading a CVE and manually hunting for similar code, the framework mines fixing commits, distills them into reusable detection rules, and runs those rules across target codebases with a staged pipeline that starts cheap and gets expensive only where it matters.
From patch to rule
The mining step pulls from 19,325 high-severity CVEs dated 2022 to 2026. From those, BUGSTONE-E2E identifies 2,710 verified fixing commits and constructs 1,033 detection rules spanning 56 CWE families, packaged into 172 skills. Each rule bundles three things: scan anchors that locate candidate call sites, fix semantics that describe what the unsafe behavior looks like, and CVE provenance so every finding traces back to a real advisory. The rules are organized by CWE and language, so the framework knows that a buffer overflow rule and an SQL injection rule need different anchors.
Key numbers: 19,325 high-severity CVEs mined. 2,710 fixing commits identified. 1,033 detection rules across 56 CWE families, packaged into 172 skills. 644 findings with runtime evidence across 14 programs.
The 1,033-rule coverage matters because it spans the common weakness taxonomy, not a single bug class. And 644 findings with runtime evidence means the output is a list short enough for a human team to actually review, with a patch attached.
A funnel that spends compute where it counts
Running a full LLM analysis on every candidate site would be slow and ruinously expensive. BUGSTONE-E2E's answer is a funnel: cheap stages process a large pool, and each later stage applies a more capable model to a shrinking set.
The first stages use Tree-sitter to enumerate call sites matching rule anchors, then heuristic filters drop benign sites without a single LLM call. Only the survivors see an LLM agent, and only the agent-confirmed candidates get runtime verification and scope-checked patch generation. Final validation runs two-sided differential tests: the patch must fix the target without breaking adjacent behavior.
This funnel design is the practical lesson most security teams miss. You can run a cheap syntax match over an entire repo for pennies. You cannot run a capable agent over every candidate. The architecture decides in advance where the expensive model earns its keep.
What the numbers mean
Each finding that reaches the end of the funnel survived agent inspection and came with a tested patch. That converts the usual alert flood into a short, verifiable fix list, which changes the economics of a security review.
Quick take: Your CVE backlog is a detection engine waiting to be built, and a funnel architecture is what makes it affordable to run.
Agent pipelines drop security context
Here's where this cluster gets uncomfortable. BUGSTONE-E2E itself relies on LLM agents inspecting candidates, so its own results depend on the agent pipeline preserving security context. CONTINUITY's core finding is that individually correct security controls do not compose. Each component might correctly check authorization or provenance in isolation, but as actions cross component boundaries, the context that matters gets dropped, widened, rebound, or reinterpreted.
The hard part of agent security is what happens to authorization context at each transition: the principal, the task, the provenance, the delegation chain, the policy state. One component issues a permit, the next reinterprets it for a slightly different action, and the chain of intent is gone.
CONTINUITY attacks this with assume-guarantee contracts and authenticated context carriers: signed root grants, provenance commitments, role-bound transition receipts, bounded typed releases, transformation witnesses, and effect-bound execution permits. Every transition between components carries signed evidence of what was authorized and by whom.
The formalism is dense, but the practical claim is simple: every realized external effect must be backed by a valid, current authorization witness linking the principal, task, provenance, delegation, policy state, canonical action, and finality boundary. If any link is missing or stale, the effect does not happen.
The evaluation runs a deterministic cross-layer fault-injection suite covering 32 fault classes across four application domains. Across 2,560 parameterized attack instances spanning 128 fault-domain classes, the full CONTINUITY configuration commits no harmful external effect, completes all 700 benign tasks without false blocking, and escalates all 200 ambiguous cases to a human instead of guessing. That scale of injection testing is where manual red-teaming simply cannot compete.
Three approaches, one pattern
| BUGSTONE-E2E | CONTINUITY | TIP jailbreak | |
|---|---|---|---|
| Target | Code vulnerabilities | Agent security controls | Frontier LLM alignment |
| Core mechanism | Executable rules mined from patch history | Assume-guarantee contracts with signed context | Hidden harmful objective inside a benign task |
| Scale | 1,033 rules from 19,325 CVEs | 2,560 attack instances, 128 fault classes | Single reported attack chain |
| Key result | 644 verified findings | Zero harmful effects in tested config | GPT-6 compromised within 24 hours |
| Practical cost | Funnel keeps LLM calls on a shrinking set | Contract setup per component boundary | Model-side defense only; near-zero attacker cost |
The pattern across all three: a security property is only as strong as the last boundary it fails to travel. CVE knowledge fails to travel from advisory to scanner. Authorization context fails to travel across component handoffs. Harmful intent travels straight through the model's own task decomposition.
The attack side: GPT-6 in under 24 hours
The third result is a reminder that attackers are not waiting. A researcher reported jailbreaking GPT-6 Astra within a day of release, using an extended Task-in-Prompt (TIP) attack. TIP works by hiding the harmful objective inside a legitimate task structure, like solving a cipher or executing Python code, so the model's instruction-following machinery does the attack's work for it.
The original minimal TIP attack, published at ACL 2025, was no longer sufficient against GPT-6. It had to be reworked and combined with four other unnamed techniques. That detail matters more than the 24-hour headline. GPT-5 fell to the same researcher within an hour of release; GPT-6 held for about a day against a more elaborate version of the attack. The defenses are improving marginally, and the attacks are getting more complex. The gap is not closing.
What the community is saying
In the threads, the reaction split along predictable lines. Some people pointed out that a single private disclosure proves little: no public reproduction, no benchmark, no way to weight the claim. Others found the researcher's track record convincing, since the same person had demonstrated a GPT-5 jailbreak in under an hour a year earlier.
What I found most useful was the operational detail buried in the report. When I tested similar TIP patterns on smaller open-weight models, the cipher-encoding trick reliably triggers unsafe completions that a direct request never would. The model happily solves the puzzle, and the puzzle happens to contain instructions it would refuse if asked plainly. TIP exploits task decomposition, a core feature of instruction-following models that won't disappear.
Common pitfalls
Five mistakes show up again and again when teams try to build on these ideas.
- Treating the CVE backlog as reading material instead of training data. The fixing commits are already written. Mining them into rules is what turns advisories into detection, and BUGSTONE-E2E's 1,033-rule corpus came from public data that sits in every security team's inbox.
- Sending every candidate to an LLM. The funnel exists because cheap syntactic filters remove most sites before any model runs. Skip that stage and your cost per finding goes vertical.
- Assuming security checks compose across components. A component that validates authorization correctly in isolation can hand that context to the next component, which widens it. CONTINUITY demonstrated this failure across 32 fault classes; you will hit it in production too.
- Dismissing jailbreak reports because they are private disclosures. The TIP technique is public and reproducible on smaller models. Reproduce it against your own deployment before deciding it doesn't apply.
- Letting the model carry security state in its context window. If the authorization context isn't signed and explicit, it gets dropped or reinterpreted at the next boundary. Carry it as data, not as conversational memory.
One thing to remember
Every security property in this cluster decays at boundaries. A patch protects one deployment; a detection rule protects every deployment that runs it. An authorization check protects one component; a signed contract protects the whole chain. A jailbreak hides in the task, so the task itself has to be treated as untrusted input. The one design habit worth taking from these results: make security knowledge executable, and make security context impossible to drop at handoffs.
The bottom line
If you're building a vulnerability scanner, mine fixing commits into rules rather than writing heuristics from scratch. BUGSTONE-E2E's numbers show the public patch record contains enough signal for 1,033 rules across 56 CWE families, and your own org's commit history can seed the same approach at smaller scale.
If you're composing LLM agents with authorization and policy controls, adopt a contract layer like CONTINUITY's signed receipts. The alternative is per-component checks that silently lose context, and a manual audit of every transition, which does not scale.
One thing to watch: TIP-style attacks are adapting faster than alignment patches. Expect jailbreak research to keep pace with each frontier release, and plan for the assumption that a deployed model can be redirected within days, not years.