PaperScope
LIVE · 2026-10-09 05:40 UTC

When AI Finds Hidden Messages, Does It Report?

William Guey, Rashik Jahangir, Pierrick Bougault, Vitor D. de Moura, Wei Zhang, José O. Gomes

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.10620 v1
Category
Submitted
2026-10-07

Abstract

When an assistant encounters a message for another AI, does it tell its user? Four fixed model-provider deployments perform simulated source tasks in 1,280 ordinary-note and 128 enhanced-note sessions. Harmless and harmful messages have matched plaintext and ROT13 versions, with no-message controls. Observers receive no decoder or decoded meaning; a requested reference code incentivizes inspection. Asking for reports increases rule-detected notifications identifying another AI as recipient by 53.1 percentage points for harmless ROT13 messages and 54.7 for harmful ones. This is a joint inspection, recognition, and notification effect; missing-response bounds are 38.3--77.3 and 36.7--78.1 points. Model-based trace checks identify eleven ordinary plaintext cases where agents interpret the message but do not notify their user. Seven encoded omissions are verified with enhanced notes; ordinary encoded omissions remain unverified. Seven simulated filename disclosures coexist with accurate review-status answers, and two answers use a planted false count. Interpretation, notification, and authorized task performance are distinct outcomes.

Comment: 2 figures, 8 tables. Data and code (v1.0.0): https://github.com/williamguey/ai-hidden-message-reporting

arXiv abs page · PDF · same-day batch