PaperScope
LIVE · 2026-09-17 05:40 UTC

Beyond Accuracy: How Procedural Traces Shift the Decision Criterion of LLM Overseers

Zihan Chen, Di Zhu, Lei Zheng, Weiling Li

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.18204 v1
Submitted
2026-09-16

Abstract

Organizations increasingly use oversight loops where one large language model (LLM) audits another's outputs alongside procedural traces of claimed steps. A common concern about such LLM-as-a-judge pipelines is that detailed traces make overseers gullible. Using signal detection theory, we audit five LLM overseers on 19 compliance tasks (4,551 analyzed judgments), varying only trace detail and evidence labeling. With disconfirming evidence always visible, error detection remains near ceiling. Instead, elaborate traces shift the decision criterion toward rejection, increasing false alarms in susceptible overseers. Without option labels, human-validated reason coding shows about 60% of false alarms cite an inability to tie evidence to its option. Labels eliminate this stated reason, yet residual rejection of correct work persists in those overseers and rises with trace detail. Procedural traces thus act as governance artifacts that shape oversight decisions. AI auditors should be evaluated by their decision criterion and false-alarm behavior, alongside accuracy.

Comment: 11 pages, 4 figures, 3 tables. Accepted at the 60th Hawaii International Conference on System Sciences (HICSS)

arXiv abs page · PDF · same-day batch