PaperScope
LIVE · 2026-10-06 05:40 UTC

Multimodal Safety Evaluation Should Measure Controllability Beyond Classification

Junhyeong Park, Hanwool Lee, DongGeon Lee, Dasol Choi, Yejin Son, Haon Park, Youngjae Yu

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.06452 v1
Category
Submitted
2026-10-05

Abstract

VLM safety is commonly evaluated through input- and output-level classification. Such classification is necessary, but it does not reveal whether a safety state is accessible or controllable inside the model. We argue that multimodal safety evaluation should therefore report a \emph{controllability profile} alongside behavioral classification, separating representation-level detectability, cross-modal specificity, intervention sensitivity, and benign-preserving selectivity. Using implicit toxicity as a stress case, we instantiate this profile on LlavaGuard and Qwen3.5 with sparse feature decompositions. LlavaGuard admits localized handles with a narrow benign-preserving intervention range and modest downstream safety gains, whereas Qwen3.5 supports strong representation-level readout but no comparable selective-control regime under the tested operators. These results show that internal readout and controllability can diverge. Future multimodal safety benchmarks should therefore report not only behavioral safety metrics, but also whether safety-relevant internal signals can be intervention-tested and controlled within a validated operating range.

arXiv abs page · PDF · same-day batch