PaperScope
LIVE · 2026-10-01 05:40 UTC

When Scientific Contradictions Are Lost in Translation

Tal Zeevi, Trey W. Jensen, Maxwell Strome

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.38621 v1
Category
Submitted
2026-09-29

Abstract

Two scientific findings can disagree without contradicting each other. Determining whether they conflict requires knowing whether they describe comparable measurements. We study how language models behave at this decision point. In a controlled task, we generate an unsatisfiable XOR constraint system and translate its constraints into scientific reports from different laboratories. One assignment satisfies more constraints, while another satisfies fewer but better matches expected biology. This creates a simple dilemma: does the model choose the assignment that best fits the constraints, or the one that better matches biological expectations? When the constraints are stated directly, GPT-5.6 Sol and Claude Opus 5 recover the best-supported assignment in 90% and 96% of cases, respectively. In scientific prose, however, the models behave differently. Claude Opus 5 often prefers the biologically expected assignment. Removing that biological preference increases recovery of the better-supported assignment from 27% to 79% (p<.001); recovery reaches 92% when the same Biology-favored record is accompanied by a formalization request and an explicit paired-design cue (p<.001). GPT-5.6 Sol is less sensitive, with neither corresponding change reaching statistical significance. These results suggest that reliable scientific verification depends not only on formal reasoning, but also on how models decide which findings should be compared and what relations they imply.

Comment: Accepted at the NeurIPS 2026 AI for Science Workshop: Verification in the Age of AI Scientists. This version is not included in the official NeurIPS proceedings

arXiv abs page · PDF · same-day batch