PaperScope
LIVE · 2026-09-18 05:40 UTC

Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models

Alizishaan Khatri, Chiquita Prabhu, Omkar Neogi

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.19472 v1
Submitted
2026-09-16

Abstract

Autonomous systems increasingly rely on Large Language Models (LLMs) yet the safety infrastructure surrounding these models introduces latency and compute overhead. This limits utility in resource-constrained, time-critical deployments. Existing external guardrail models remain blind to the model's internal workings, creating a fundamental assurance gap. We ask: does the model already know when the content is harmful? We extract activations from LLaMA-3.1-8B and train lightweight MLP classifier probes (12.6M parameters) to detect harmful prompts. Evaluated on WildJailbreak, Beavertails, and AEGIS 2.0, our probes achieve F1 scores of 99%, 83%, and 84%, respectively competitive with 1000x larger guard models while cutting latency and compute costs.

Journal: IEEE DSN-W 2026, pp. 48-52

arXiv abs page · PDF · same-day batch