PaperScope
LIVE · 2026-09-15 05:40 UTC

Measuring the Creativity of Frontier LLMs in Automated Research

Yiheng Zhao, Mengzhuo Chen, Chengming Hu, Pengyi Liao, Yiran Pang

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.14057 v1
Category
Submitted
2026-09-12

Abstract

Frontier LLMs are increasingly capable of conducting automated research, yet their creativity in this setting has not been systematically evaluated. In this paper, we propose a set of metrics to evaluate creativity along the two dimensions of valueness and novelty. Valueness assesses whether each proposed idea is useful, while novelty is evaluated from three perspectives: whether the same idea has appeared before (Exact-Match P-Novelty), whether a previously unexplored variable or variable combination is explored (Variable-level P-Novelty), and whether the idea directly follows retrieved external knowledge or departs from it (H-Novelty). Our evaluation shows that the models achieve relatively similar scores on most creativity metrics, but differ substantially in Variable-level P-Novelty, which reflects the breadth of research-space exploration. Further correlation and idea-level performance analyses show that Variable-level P-Novelty is the creativity dimension most strongly associated with research performance.

arXiv abs page · PDF · same-day batch