PaperScope
LIVE · 2026-09-09 05:40 UTC

We're Cooked! - Probing LLM Political Alignment Via Conflict-Framed Recipe Translation

Svetlana Gorovaia, Angelica Henestrosa, Ivan P. Yamshchikov

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.07568 v1
Category
Submitted
2026-09-07

Abstract

Large language models (LLMs) are increasingly deployed for translation tasks, yet their implicit political positioning in such contexts remains understudied. We ask whether a single politically charged framing term, such as aggressor, enemy, neighbour, or coloniser is sufficient to trigger implicit political alignment in an otherwise apolitical task. We present a fully crossed factorial study in which eight models spanning Western, Chinese, and European origins are prompted to translate culturally attributed recipes into a target language left deliberately unspecified. Across 17 languages, four framing conditions, eight models, and 15,680 responses, we find that models do not simply decline or ask for clarification but resolve the ambiguity. Language resolution and reasoning behavior cluster meaningfully along model families: Western models hedge and deflect with vague justifications, Chinese models resolve conflicts silently, and Mistral Large emerges as a distinct profile combining high compliance with conflict-grounded reasoning. Sensitivity to framing terms is consistent across models: even subtle framing variation is sufficient to modulate behavior. Our findings urge caution when deploying LLMs for translation in conflict-adjacent contexts, where implicit political judgments may be made without any signal to the user.

arXiv abs page · PDF · same-day batch