PaperScope
LIVE · 2026-09-03 05:40 UTC

Sleight of Word Benchmark: Can Language Models Notice If Their Own Output Was Tampered With?

Alberto Cetoli

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2608.29921 v1
Category
Submitted
2026-08-30

Abstract

The output of a Language Model can be tampered with \emph{while} the model is writing it. A simple test can thus be constructed by evaluating the model's perception of this external perturbation. In this spirit, a simple benchmark is built in which a single word is consistently substituted with another in the generation process. We call this method \emph{Sleight of Word}. Two distinct axes are measured: metrics that relate to the model's surprise, as well as an evaluation of the textual reaction for 19 different open-weight language models.

Comment: Accepted at INLG 2026

arXiv abs page · PDF · same-day batch