PaperScope
LIVE · 2026-09-29 05:40 UTC

Artificial intelligences and human scientists exhibit complementary strengths in theory building

Ke Li, Spyros I. Zoumpoulis, Phanish Puranam, Philip Parker, Matthew Eshbaugh-Soha, Izzy Gainsburg, Michael Gilead, Igor Grossmann, Britt Hadar, Yoel Inbar, Almog Simchon, Robb Willer, Rui Ai, Ruicheng Ao, Gavin J. Bala, Matthew Bidwell, Shuang Cai, Kai Chang, Skyler Y. Chen, Cory J. Clark, Irmak Dai, Abhinandan Dalal, Connor Douglas, Alexis Du, Zhehang Du, Leyun Feng, Isabel Fernandez-Mateo, Linnea Gandhi, Cyrille Grumbach, Anmol Gupta, Vansh Gupta, Maria Hademer, Jay H. Hardy, Chen Kai Huang, Jacob Xiangyu Jin, Ufuk Keskin, Na Hyun Kim, Mert Kobaş, Byounghoon Koh, Gabrielle Lamont-Dobbin, Gregory Lanzalotto, Sun Young Lee, Dingzhe Leng, Chenjun Li, Weiyuan Li, Zeyuan Li, Zhongyuan Liang, Ning Liu, Peihong Liu, Yuhan Liu, Jiuyao Lu, Wanteng Ma, Nicolas Martinet, Natnael Mulat, Christina A. Nguyen, Khai Nguyen, Quang Minh Nguyen, Naja Pape, Chanwoo Park, Stefanos Poulidis, Jeffrey Sanchez-Burks, Michael Schaerer, Isabelle Solal, Yanbo Song, Junghyo Sun, Qingyao Sun, Rui Sun, Roderick Swaab, Kevin Tan, Dequn Teng, Michelle A. Vaccaro, Robin Vigerbaeck, Xiaomeng Wang, Randol H. Yao, Duygu Yilmaz, Shun Yiu, Ecem Yucesoy, Allen Zang, Ruijia Zhang, Xilan Zhang, Yichi Zhang, Zhanhao Zhang, Eric Luis Uhlmann

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.32562 v1
Category
Submitted
2026-09-26

Abstract

We investigate the effectiveness of artificial intelligences (AI)-specifically large language models (LLMs)-relative to human scientists at high-level cognitive tasks in social science such as theory formulation, predictions of novel empirical results, and theory revision in response to new evidence. The research domain was academic discourse regarding gender and race inequality. Our findings, comparing 25 LLMs with 13 senior researchers and 60 doctoral scholars, reveal that the AIs outperformed most humans individually on most of the present tasks, while human theories were more diverse and exhibited greater gains in predictive accuracy from aggregation. AI-generated theories were more extensively elaborated, involving additional theoretical paths and latent variables, and were rated as higher quality than human theories by independent raters blinded to source. However, this theoretical complexity was in part ornamental, in that it was not associated with more accurate predictions about empirical patterns in data; in contrast, human scientists achieved greater predictive efficiency with simpler theories. The AIs were significantly more likely than human scientists to revise their theories to incorporate new evidence; human scientists updated their beliefs in a selective way that is sensitive to prior prediction errors. We speculate that the superior processing capacity of artificial intelligences makes them especially well-suited to tasks requiring grappling with complexity, but that the greater diversity of human ideas is essential to wise crowds and collective creativity.

arXiv abs page · PDF · same-day batch