PaperScope
LIVE · 2026-09-15 05:40 UTC

ForeSight: Enhancing Risk Monitoring via Early Safety Signal Distillation

Hanling Wang, Chenlong Wei, Ling Xu, Hanyan Niu, Qi Cao, Shizhou Huang, Yang Yang, Xiaohui Zhu, Yao Zhu

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.13737 v1
Category
Submitted
2026-09-12

Abstract

As large language models (LLMs) are increasingly deployed, the generation of harmful content has become a critical safety concern. Existing safeguards operate at the input, output, or streaming-generation stages, while early-risk methods that rely on surface tokens or output logits may suffer from weak initial signals, and internals-based detectors using dense representations may retain highly entangled and redundant safety-irrelevant information. It therefore remains unclear whether the earliest post-generation hidden states already contain reliable signals about final-response harmfulness. To address this gap, we propose ForeSight, a first-token output-risk forecasting framework that distills weak and redundant early safety signals into compact, layer-aware risk representations. Experiments on five safety benchmarks and two target models demonstrate that ForeSight achieves superior and efficient early-risk forecasting while relying solely on first-token hidden states. The code is available at: https://github.com/Scabbards1500/Foresight

Comment: 17 pages, 12 figures. Accepted to Findings of EMNLP 2026

arXiv abs page · PDF · same-day batch