Appearance
The safety gap no one measures
Every safety intervention in this cluster has a failure mode that only shows up in production. Unlearning methods help a model forget copyrighted text, but the knowledge creeps back after you quantize to 4-bit. Safety fine-tuning raises refusal on harmful prompts, but the same training updates turn a helpful assistant into a machine that refuses three out of four safe queries. Persona steering protects against malicious fine-tuning, until you realize the protection never actually lived in the weights.
Three recent papers attack these failures directly. FOM-UL makes unlearning survive post-training quantization. Boundary-aware self-distillation scopes safety refusal to a subset of a topic, not the whole topic. A new analysis of Preventative Steering shows that adversarial defense works only while it keeps adapting. Underneath all of it sits a governance story: Paul Christiano just joined the OpenAI Foundation Board and its Safety and Security Committee.
The through-line is simple. Safety is not a property you bake into a model once. It is a boundary you scope, measure, and maintain. Everything below is about how to do that without wrecking the model.
Forgetting only what matters: layer-selective unlearning
Machine unlearning is the escape hatch for content you cannot ship: copyrighted books, leaked chat logs, dangerous instructions. Full retraining is too expensive, so you patch the weights. But most unlearning methods apply broad or fixed parameter updates, and two things go wrong. The updates degrade general utility. And when you quantize the model for cheaper hardware, small diffuse changes get rounded away, so the forgotten knowledge partially re-emerges.
FOM-UL (Forgetting Only What Matters via Unlearning Layers) takes a different route. Instead of updating the whole model, it scores each transformer layer with a forget-to-retain significance measure. Layers that strongly influence the forget set but barely matter for the retain set get updated. Everything else stays untouched. The change is concentrated where it works.
On TOFU, KnowUnDo, and MUSE-style evaluations, FOM-UL cuts residual memorization compared with GA, NPO, KLD, SURE, ReLearn, and LUNAR-based baselines, and retain-set utility stays close to the vanilla model. Under 8-bit and 4-bit post-training quantization, the suppression holds stronger than competing methods, and adversarial prompts recover less forgotten content.
The 8-bit and 4-bit detail is the deployment story. That is the range where a 7B model fits on a single consumer GPU or runs on edge hardware, and it is exactly where low-bit rounding erases small safety updates. FOM-UL's layer concentration survives rounding because the update is big where it matters.
One caveat, stated plainly in the paper: FOM-UL offers no formal guarantee of erasure. It is an empirical method. If your compliance team needs a legal claim that content is gone, no current unlearning technique gets you there, layer-selective or not.
Refusing a subset, not a whole topic
The second paper, Safety for Whom?, starts from a deployment fact: real products rarely need topic-level refusal. A civics tutor and a public-sector assistant can share the same base model and want opposite behavior on politics. Both should answer factual questions about an election. Only one has to refuse a request to write targeted political manipulation. A topic-level guard cannot express that split.
LlamaGuard-3 shows why. It defines the election category as "factually incorrect information about electoral systems and processes." Persuasion and manipulation fall outside that definition, and a topic-level refusal triggered on the category would sweep up the factual prompts a deployment must keep answering. The guard does the wrong thing in both directions at once.
The paper formalizes the problem as a topic universe with a target-harmful subset. The intended policy is a sharp step: refuse inside the subset, answer everywhere else. A trained model never learns a sharp step. It learns a refusal probability, and cross-entropy training that raises refusal inside the harmful subset can push refusal outward into the benign complement. The fix is to train and measure on the boundary itself: pairs of prompts that share a topic anchor and differ only in intent, one to refuse and one to answer. The paper uses political persuasion as the testbed on Qwen3-8B, because manipulative persuasion causes real harm while factual political information stays legitimate.
Quick Take: Safety tuning that reports only refusal rates will fool you, because every intervention shifts the boundary between refusal and compliance in both directions.
Where self-generated safety data breaks
The natural way to build safety training data is self-generation: steer the target model toward refusal on each harmful prompt, and keep the traces a guard model verifies as genuine refusals. This is the recipe behind methods like ThinkSafe. Framing the task as a boundary instead of a topic exposes three weaknesses in that standard pipeline.
The first is a coverage gap. A single steering attempt does not always produce an accepted refusal, and the naive pipeline silently drops those prompts. In the audited pool, single-shot generation dropped 19.88% of prompts, 8,009 of them. The dropped prompts are probably the hardest examples, the ones you least want to lose. The fix is escalating retry: resample the same prompt through progressively stronger steering until a verified refusal appears. That brings residual failures down to 0.20%, or 79 prompts, leaving 40,293 harmful training prompts where the naive pipeline would have thrown thousands away.
The second is downside reactions. Safety tuning produces false refusals on benign prompts that look superficially dangerous. The paper compensates with in-distribution benign data, including 11,955 verified surface-dangerous benign prompts across 18 semantic types. The model sees safe prompts with dangerous-looking wording during training, not just at evaluation, so it learns the difference.
The third is measurement. Ordinary harmful/benign splits do not measure the shape of the boundary at all. A model can improve its harmful-refusal rate simply by expanding refusal into nearby permissible prompts, and a topic-level metric will call that an improvement. Held-out harmful-benign pairs, 1,539 per side, let you measure both sides of the boundary directly.
The numbers that hide a blunt refusal machine
Training on political refusal data works in the obvious sense. On Qwen3-8B, the escalated-coverage model raises in-distribution political refusal from 9.47% to 84.75%, so the model that almost never refused now refuses consistently on political persuasion. The effect also transfers: the mean unsafe-response rate across HarmBench, StrongREJECT, and WildJailbreak, scored by LlamaGuard-3, falls from 26.26% to 0.14%.
Reported alone, those numbers look like a clean win. They are not. At the same checkpoint, over-refusal on XSTest rises from 2.00% to 74.00%. The configuration with the lowest harmful-response rate refuses nearly three quarters of plainly safe prompts. In production that means blocked support requests, broken workflows, and users who learn to phrase around the guardrail. The checkpoint is a blunt refusal machine. I have made this exact mistake: shipping a model that scored beautifully on refusal benchmarks and then watching it decline a benign question about how elections work.
The data composition decisions pull over-refusal back down without giving up the safety gain. Replacing externally adopted compliance responses with verified responses generated by the target model itself lowers XSTest over-refusal from 15.20% to 5.20% at a modest harmfulness cost. The boundary pairs do the most precise work: adding benign boundary data cuts over-refusal on the comply-worthy side of the held-out pairs from 32.94% to 4.16%, while refusal on the harmful side drops only from 91.88% to 87.72%. Most of the false refusals near the boundary disappear, almost all the genuine refusals survive, and the recall cost is small and measurable.
84.75%: in-distribution political refusal on Qwen3-8B after escalated-coverage training, up from 9.47%. 0.14%: mean unsafe-response rate across HarmBench, StrongREJECT, and WildJailbreak in the strongest configuration. 74.00%: over-refusal on XSTest at that same checkpoint, up from 2.00%. 19.88% to 0.20%: harmful prompts lost by single-shot refusal generation, 8,009 prompts, cut to 79 by escalating retry. 32.94% to 4.16%: over-refusal on the comply-worthy side of held-out boundary pairs after adding benign boundary data.
Active adaptation beats static defense
The third paper asks why Preventative Steering works at all. The defense injects undesirable-trait persona vectors during fine-tuning and removes them at evaluation time. Looking at the temporal optimization dynamics, the protection emerges from an early compensatory adaptation phase, then a steady-state phase where the corrective signal decays. In parameter space, attention output projections are the dominant residual-write route for the defensive updates.
Then the paper tries to break its own defense. Intervention Delta Preservation keeps the weight offset in place, and IDP Continuation reinjects it later. Neither maintains protection. That is the key result: preventative steering relies on active adaptation, and preserving the weights that carried the defense does nothing, because the defense never lived only in the weights.
That insight leads to Progressive Intensity Scheduling (PIS). Start with a moderate injection strength, then increase it after static-strength alignment begins to decay. On Qwen2.5 and Gemma-3, PIS improves safety robustness over static-strength steering while reducing harmful trait expression.
The three interventions side by side:
| Approach | What it modifies | How it decides | Headline result | Caveat |
|---|---|---|---|---|
| FOM-UL | Selected transformer layers | Forget-to-retain significance score | Beats GA, NPO, KLD, SURE, ReLearn, LUNAR on TOFU, KnowUnDo, MUSE | Empirical only, no formal erasure |
| Boundary-aware self-distillation | Full model through curated data | Harmful-benign pairs near the refusal boundary | Over-refusal on boundary pairs drops from 32.94% to 4.16% | Needs a per-deployment data pipeline |
| Preventative Steering + PIS | Persona vectors during fine-tuning | Injection strength scheduled by decay detection | Better robustness than static steering on Qwen2.5 and Gemma-3 | Requires fine-tuning access and active monitoring |
Governance catches up
None of this lands without an organizational decision to treat safety limits as scoped and accountable. Paul Christiano joining the OpenAI Foundation Board and its Safety and Security Committee is a signal that alignment decisions are moving from the training run to the boardroom. His background combines AI alignment, safety, and standards work, which is a rare mix for a foundation board seat.
This matters for your deployment concretely. Board-level attention changes the incentive structure. If a refusal boundary is written into a documented deployment policy, then someone is accountable for what gets refused and what gets answered. The papers above give you instruments to measure the boundary. Governance gives you a reason to measure it.
Common pitfalls
Five things trip up teams when they try to apply this work.
Measuring only the harmful side. If your eval suite reports refusal on harmful prompts but not over-refusal on safe prompts, you will ship a model that refuses 74% of XSTest prompts and call it safer. Report both axes from the same checkpoint, in the same run.
Trusting unlearning through quantization. If you unlearn with broad diffuse updates and then quantize to 4-bit, rounding can partially resurrect the forbidden knowledge. Concentrate updates on high-significance layers, and verify suppression after quantization, not before.
Assuming safety data coverage. When I built a self-generated refusal dataset, I found that a single steering pass silently dropped nearly one in five prompts, and the failures were the hard cases. Use escalating retry until the guard model accepts the refusal. Dropped prompts do not come back on their own.
Confusing static weights with active defense. If your defense relies on a persona vector injected during fine-tuning, the protection lives in the adaptation process, not the final weights. Preserving the weight offset does not preserve the defense. Monitor for decay and schedule reinforcement.
Treating topic-level guards as boundaries. The civics tutor and the public-sector assistant genuinely need different behavior on the same topic. If your guard cannot express that split, it will over-refuse in one deployment and under-refuse in the other.
One thing to remember: safety is measurable, and the measurement must include both sides of the intended boundary plus the deployment conditions after training. A number that only tells you the harmful-refusal rate is telling you half the story.
Takeaways for deployment teams
If you are shipping a model that must forget copyrighted or sensitive content, adopt a layer-selective method like FOM-UL and verify suppression after quantization, because broad updates will not survive 4-bit rounding and the content will creep back.
If you are fine-tuning for a specific deployment policy, scope refusal to a harmful subset instead of a whole topic, and build boundary pairs into both training and evaluation, because topic-level guards will over-refuse in one product and under-refuse in another.
If you are defending against malicious fine-tuning, use an adaptive schedule like PIS and plan to monitor decay continuously, because static defenses live in the training dynamics, not in the weights, and preserving the weights does not preserve the protection.