PaperScope
LIVE · 2026-09-09 05:40 UTC

Do Reasoning Representations Help Humans Evaluate LLM Outputs?

Jaewoo Lim, Sungbok Shin, Sanghyun Hong

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.09038 v1
Category
Submitted
2026-09-08

Abstract

Reasoning representations are increasingly used as explanations for large language model outputs. Yet they are typically evaluated with model-centric criteria, such as answer accuracy and faithfulness, leaving it unclear whether they help people evaluate model responses. In this work, we study reasoning representations as human-facing interfaces rather than proxies for model reasoning ability. We conduct a controlled human study of six reasoning formats across tasks of varying complexity, supported by a web-based framework that randomizes task domains, problem instances, and representation order. The study collects fine-grained judgments of structural understanding, error detection and localization, and trust calibration. Our study shows a mismatch between perceived preference and support for human evaluation. Participants prefer planning- and decomposition-based representations, but simpler chain-of-thought traces better support verification, trust, and interpretability. Preferred representations also introduce calibration risks, with more false alarms on correct traces and high trust despite low willingness to verify.

Comment: 19 pages. EMNLP 2026 (Findings)

arXiv abs page · PDF · same-day batch