PaperScope
LIVE · 2026-09-09 05:40 UTC

MARBO: Relational Belief Grounding for LLM Agents in Social Deduction Games

Hwang Yechan, Bae Sangjun, Kim Jeongmo, Bang Sangwoo, Han Seungyul

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.06563 v1
Category
Submitted
2026-09-06

Abstract

Social deduction games (SDGs) require agents to reason under partial observability by maintaining relational beliefs about hidden roles and team alignments. While recent LLM-agent approaches improve gameplay through prompting and preference optimization, they often optimize actions and in-game speech without explicitly grounding them in such beliefs. This frequently leads to strategically inconsistent behavior, especially for compact LLM agents. We introduce Multi-Agent Relational Belief Optimization (MARBO), a belief-grounded preference optimization framework that leverages relational beliefs to guide strategic decisions and in-game speech. MARBO provides preference feedback only when behaviors are supported by reliable relational beliefs and lead to strategically favorable social outcomes, encouraging more consistent learning under uncertainty. Experiments on representative SDGs show that MARBO enables compact LLM agents to consistently outperform existing baselines. The Code is available on https://github.com/PleaseTakemeAway/MARBO.

Comment: 9 pages, accepted to EMNLP 2026

arXiv abs page · PDF · same-day batch