PaperScope
LIVE · 2026-09-25 05:40 UTC

Artificial Societies Benchmark: A Validation Framework for Synthetic Research

Edoardo Chidichimo, Min Jun Jung, Felix P. S. Wallis, James K. He

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.30030 v1
Category
Submitted
2026-09-24

Abstract

A synthetic survey can reproduce the average answer while misrepresenting how people differ, how their answers relate to one another, or how they respond to changes in conditions. We introduce the Artificial Societies Benchmark to help researchers assess whether synthetic populations support their intended analyses. The framework combines eleven tests across internal, construct, and external validity, drawing on twenty human sources and comparing nine language models. It connects each research use to the evidence it requires and tests how results change with the information we supply about respondents. Importantly, strong performance in one domain does not establish fidelity in the others. Models often answer too consistently, compress response scales, and alter relationships between traits whilst richer profiles improve prediction for some models and worsen it for others. The resulting scorecard helps researchers identify which aspects of a synthetic population can support their analysis and where researchers need further human evidence.

Comment: 36 pages, 9 figures, 9 tables

arXiv abs page · PDF · same-day batch