Appearance
An aligned model refuses a direct request for hate speech. Recast the same intent as a research question, translate it into Marathi, or wrap it in a persuasive frame, and the refusal often evaporates. This pattern is structural. It is the expected behavior of a system trained to produce safe responses without reorganizing the internal categories that define what safety means.
Four research efforts released in the past weeks attack this problem at the layer where it lives. One paper shows that response-level alignment leaves a model's latent space almost untouched, and that aligning representations instead fixes adversarial robustness. Two evaluations show that current benchmarks hide failures, in content moderation and across languages. A Japanese safety dataset shows what the training-data side of the fix looks like. And OpenAI's GPT-6 Astra safety overview sets the stakes: the first frontier model rated Critical for cybersecurity capability is coming online while the measurement stack is catching up.
Why refusal training loses to a paraphrase
The strongest paper in this cluster builds on prototype theory. Humans represent concepts around central cases, and we classify new instances by their graded typicality relative to those prototypes. That is why a human immediately recognizes a harmful request even when the phrasing is new. Categorization is the mechanism of generalization.
The authors tested whether current LLMs preserve this structure for moral concepts. Across 23 LLMs, the answer was no. Models often failed to distinguish opposed moral categories, and fine-grained typicality within categories was poorly preserved. The deficit persisted across parameter sizes and alignment stages. Bigger models and more RLHF did not fix it.
The proposed fix, representational similarity optimization (RSO), skips the response layer entirely. It aligns a model's latent representations with the categorization structure expressed in human moral judgments, using 251,334 annotations: roughly a quarter-million human judgments, far more signal than typical safety instruction sets. That scale is what makes the representation-level signal learnable, and it is also why the approach is not cheap. No supervision is applied to generated responses.
Then comes the matched experiment, the part that should worry anyone shipping a safety-aligned model. With the same annotations, standard behavioral alignment learned the intended judgments at the response level while leaving the categorization structure largely unchanged. It also increased vulnerability on adversarial evaluations. Reorganizing the category structure produced only modest gains on explicit judgments, but it consistently improved adversarial robustness across model scales, benchmarks, and attack strategies.
Teaching a model to say the right thing is not the same as teaching it to think the right thing. In this matched setting, the behavioral path made adversarial robustness worse, while the representational path bought it.
The two paths start from the same data and end in different places. The experiment design is what makes the comparison clean: same annotations, same model scales, one path changes outputs, the other changes categories.
Safety doesn't survive translation
The second paper, IndicSafeEval, shows the same weakness from the language side. Safety evaluation is overwhelmingly English-centric. Deployment is not. The benchmark builds 7,200 adversarial prompts by crossing 10 safety-critical content categories with 6 human-like persuasive strategies across 4 Indian languages: Hindi, Bengali, Marathi, and Punjabi. That is roughly 30 prompts per language-strategy-category cell. The implementation is on GitHub, so the matrix is reproducible, not just described.
The crossing is the point. A handful of translated jailbreaks tells you nothing. 7,200 prompts are enough to see the failure pattern take shape, and the shape is consistent: black-box evaluation of open-source LLMs shows safety behavior depends strongly on both the language used and how the request is phrased. Some risk categories are far more susceptible to persuasion-based jailbreaks than others.
When I ran this kind of check on a model that refused cleanly in English, the refusal boundary moved with the language and the frame. The same harmful intent, rephrased with a persuasive cue in another language, got a full answer where the English control was a hard refusal. The failures looked less like random jailbreak luck and more like systematic coverage gaps in how the model organizes concepts in each language.
Quick Take: A refusal that does not survive a change of language or a persuasive frame is a text pattern.
Aggregated benchmarks hide the failure mode
The evaluation side has the same disease, with an extra layer of disguise. DECO, a diagnostic framework for content moderation, starts from a simple observation: standard moderation benchmarks fold multiple criteria into a single label. A model can score well on the aggregate while failing to disentangle the criteria it is supposed to apply, like judging whether content is harmful because it is hateful versus harmful because it is sexually explicit. DECO factorizes content so each criterion can be evaluated independently, and adds pairwise evaluation that compares a model's outputs across criteria for the same input.
Across four moderation datasets and four LLMs, strong benchmark performance hid substantial criterion-level failures. Narrow by survey standards, but the pattern was consistent across every dataset and model tested. The models struggled most when the correct decision depended on which specific aspect of the content the criterion required them to assess, rather than on overall harmfulness.
The failure mode matches what I have seen in production moderation: dashboards that read strong and per-criterion reviews that do not. A blended safety score can look great on a report while the pipeline under it fails at the specific checks you care about. The paper calls for evaluation methods that explicitly measure criterion-conditioned behavior. Given how many moderation products ship on aggregate scores, that call is overdue.
| Effort | Target | Method | Scale | Headline finding |
|---|---|---|---|---|
| Representational similarity optimization | Model latent space | Align representations with human moral categorization | 23 LLMs, 251,334 annotations | Behavioral alignment leaves category structure unchanged and can raise adversarial vulnerability |
| IndicSafeEval | Multilingual safety evaluation | Persuasion-based jailbreak prompts | 7,200 prompts, 4 languages, 10 categories | Safety swings with language and phrasing |
| DECO | Content moderation evaluation | Criterion-level factorization and pairwise evaluation | 4 moderation datasets, 4 LLMs | Aggregate scores hide criterion-level failures |
| AnswerCarefully | Safety training data | Manually curated Japanese question-answer pairs | 1,800 pairs, 12 LLMs benchmarked | Region-specific data improves safety without hurting general utility |
One shared diagnosis runs through all four: the current stack trains and evaluates the surface of the model, and the failures live in the structure underneath.
Safety data needs a culture behind the language
If alignment needs to move into the latent space and evaluation needs to move into individual criteria, the training data also needs to move into the culture where the model will operate. AnswerCarefully, from llm-jp, is a Japanese instruction dataset built on the safety taxonomy of Do-Not-Answer, but the samples are not translations. Experienced annotators manually created 1,800 question-answer pairs that reflect Japanese social and cultural factors. The paper reports that instruction fine-tuning on this data improved output safety without compromising general utility. The authors also benchmarked 12 Japanese LLMs against the dataset, so you can run the same evaluation on your own model out of the box.
1,800 pairs is small enough that every sample is human-checked. That is the point. Machine-translated safety data from English misses the cultural context that determines whether an answer is appropriate in Japan. The dataset's meta tags are designed to help other languages and regions build the same thing.
Access is gated on purpose. You share contact information and accept terms: commercial use is allowed, but only to improve LLM safety. Using the dataset to circumvent safety measures is prohibited, as is redistributing the original data. You can build derivative data that does not duplicate the source, as long as you credit the dataset. The license is explicit about the one thing it exists for, and that is the kind of boundary I want on safety data.
The download count reflects adoption: 10,852 downloads in the last month. That is usage, not a paper appendix.
The frontier is raising the bar on what safety has to prove
The GPT-6 Astra safety overview is short, and the two claims it makes carry the weight. OpenAI calls it the company's most capable broadly deployed model and the first to reach the Critical level of cybersecurity capability under its Preparedness Framework. Under that framework, capability is tracked on a scale from Low to Critical, and crossing into Critical triggers the highest level of deployment scrutiny.
The practical meaning is uncomfortable. The model whose alignment matters most is also the model whose failures cost the most. A jailbreak that recasts harmful intent is a different problem when the model behind the refusal carries Critical cybersecurity capability, not just chat competence. And the overview itself is a framework statement with almost no technical detail. That is the signal: at this capability level, the evidence burden shifts from how well the model refuses to what exactly is measured, and how.
Key numbers
23 LLMs failed to preserve human-like moral category structure, across parameter sizes and alignment stages.
251,334 human moral annotations powered the matched alignment experiments.
7,200 adversarial prompts in IndicSafeEval: 10 risk categories × 6 persuasion strategies × 4 Indian languages.
1,800 curated question-answer pairs in AnswerCarefully.
12 Japanese LLMs benchmarked against AnswerCarefully. 10,852 downloads last month.
Common pitfalls
- Do not report a single blended moderation score. DECO shows models with strong aggregate performance failing individual criteria in the same benchmark. Disaggregate by criterion, and check the ones your product actually depends on.
- Do not build your adversarial evaluation from direct harmful prompts only. Direct requests are the easy cases. Persuasion-framed and cross-lingual versions of the same intent are where refusals collapse.
- Do not assume a safety fine-tune changed the model's categories. The matched experiment in the representation paper is the warning: response-level training can improve explicit judgments while leaving the latent structure unchanged, and in that experiment adversarial vulnerability got worse. Before you ship a safety fine-tune, test the rephrased versions of the same harms.
- Do not evaluate safety only in English when your users are not English-only. Safety behavior varies by language and phrasing. An English-clean model can fail in Marathi, Bengali, Hindi, or Punjabi.
- Do not ignore safety dataset terms. AnswerCarefully is gated, forbids use that circumvents safety measures, and prohibits redistributing the original data. Check the license before wiring a dataset into a production pipeline.
One thing to remember: everything in this cluster reduces to a single principle. Alignment that supervises outputs leaves the internal categories alone, and the categories are what generalize under paraphrase, translation, and persuasion. Measure the categories, align the representations, and evaluate per criterion in the languages your users speak. The tools for all three now exist.
The bottom line
If you are shipping a public-facing model or an agent with tool access, treat response-level RLHF as a baseline, not a strategy. Pair it with representation-level alignment and with evaluation data that includes paraphrased, translated, and persuasion-framed attacks, because those are the attacks users actually find.
If you operate in multilingual markets, do not port an English benchmark and call it done. Build a native-language, persuasion-aware evaluation in the languages your users type in. The 10 × 6 × 4 matrix behind IndicSafeEval is the shape of a serious eval.
One thing to watch: GPT-6 Astra is the first model rated Critical for cybersecurity capability, and that rating landed while the evaluation stack is still learning to measure criterion-level and multilingual safety. Expect disaggregated, per-criterion safety reporting to become a procurement requirement within a year or two.