PaperScope
LIVE · 2026-09-29 05:40 UTC

AnchorRep: Defending LLMs Against Cross-Model Adversarial Transfer via Representation Repulsion

Gal Wertheizer, Rom Himelstein, Tomer Peretz, Avi Mendelson

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.32602 v1
Category
Submitted
2026-09-26

Abstract

Adversarial attacks optimized on a single open-weight LLM can transfer to and jailbreak architecturally different models, allowing an attacker with white-box access to one model to compromise independently deployed systems. This creates a shared vulnerability across models, yet existing defenses are not designed for this cross-model threat. We find that cross-model transfer aligns with shared internal representation geometry, making it a natural defense target. AnchorRep targets this geometry directly with a lightweight LoRA adapter that pushes the defended model's internal representations of harmful prompts away from those of a frozen anchor model on the same prompts. Training uses a small set of harmful prompts and no adversarial examples. Across five models and four architectural families, AnchorRep reduces cross-model attack success rate to <=1.1% on 2,000 transferred attacks (0% on two), including the largest drop on Mistral (36% -> 1.1%). Existing defenses can reduce transfer, but only at high cost either inducing up to 77% degenerate benign output or increasing over-refusal by up to 18%. Because such degenerate benign outputs are not captured by standard refusal-based metrics, we introduce the Benign Garble Rate to quantify them. Our results suggest that cross-model robustness can be achieved by shaping representation geometry, without requiring attack-specific training

Comment: NeurIPS 2026; Code, configs and logs: https://github.com/galwert/AnchorRep

arXiv abs page · PDF · same-day batch