PaperScope
LIVE · 2026-10-01 05:40 UTC

The Concrete-Arbitrary Gap: Kinship Reasoning in LLMs Is Not Indifferent to Presentation

Thomas Pashby

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.39913 v1
Category
Submitted
2026-09-30

Abstract

We test whether large language models solve formally matched kinship problems equally well when relations are expressed in familiar vocabulary or by explicitly defined nonce predicates. Across 500 paired graphs, concrete accuracy exceeds arbitrary accuracy by 35.6 percentage points in local Qwen3.8-27B, 26.6 in Gemma 4 26B-A4B, 12.0 in Gemma 4 31B, and 5.4 in Qwen3.8-Max. All four paired gaps are statistically resolved. Reasoning budgets and prompt-language interventions can substantially reduce the difference, showing that it is modifiable rather than a fixed incapacity. The minimal conclusion is behavioral: on these tasks, the models' manifested relational competence is not indifferent to presentation. Explicit definitions provide the formal relations but do not make nonce predicates as usable as familiar vocabulary embedded in learned linguistic associations.

arXiv abs page · PDF · same-day batch