PaperScope
LIVE · 2026-09-29 05:40 UTC

Beyond Correctness: Evaluating Semantic Knowledge in Cross-Table Transfer

Seokyong Sheem, Hochang Lee, Suyeong Lee, Daekyum Kim

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.34098 v1
Category
Submitted
2026-09-28

Abstract

Semantic knowledge is increasingly used to bridge heterogeneous schemas in tabular learning, but how much does that knowledge actually improve prediction? Studies in tabular learning commonly answer this question through semantic ablations that modify or suppress the supplied semantic knowledge. We show that these ablations can lead to misleading conclusions about predictive benefit: poor performance under altered semantics may be taken as evidence that the intended knowledge is beneficial. Across real and controlled experiments, altering semantic content can produce large performance differences even when the model gains little predictive benefit from having that semantic knowledge in the first place. To separate these effects, we distinguish two quantities: content sensitivity and predictive utility. Content sensitivity measures the change in performance when semantic content is altered, whereas predictive utility measures the benefit of the intended semantic knowledge relative to a suitable reference without that knowledge. This distinction motivates an evaluation framework in which the control is chosen according to the question being asked: altered controls assess sensitivity to semantic content, whereas claims that semantic knowledge improves prediction require a suitable reference. Even then, predictive utility is not fixed; it varies across suitable references and decreases when the reference can more easily recover the tested knowledge from other inputs or labeled examples. In a bounded audit of 25 semantic-ablation comparisons across nine studies, only one of 18 explicit predictive-utility claims is paired with a control that clearly isolates the tested semantic contribution. Together, these findings motivate a simple evaluation principle: semantic-ablation controls should be chosen and interpreted according to the question they are intended to answer.

Comment: 40 pages. Seokyong Sheem and Hochang Lee contributed equally. Corresponding author: Daekyum Kim. Code and reproducibility artifacts: https://github.com/mintlabkorea/beyond-correctness

arXiv abs page · PDF · same-day batch