PaperScope
LIVE · 2026-10-06 05:40 UTC

Spec2Game: Can LLMs Generate Complete Playable Games from Detailed Specifications?

Yixue Cai, Yuzhe Zhao, Hanxiang Chao, Qingsen Ma, Ziheng Xiong, Jinhu Qi, Irwin King

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.04253 v1
Category
Submitted
2026-10-03

Abstract

Generating an executable program does not necessarily mean that it correctly implements the behavioral requirements specified in natural language. To evaluate large language models' ability to realize detailed specifications as complete interactive programs, we introduce Spec2Game, a benchmark that requires models to generate complete Pygame projects from detailed natural-language game specifications. Spec2Game comprises 15 game families and 150 task instances, with one canonical task and nine controlled rule variants per family, spanning three levels of implementation complexity. Using source-code, runtime, and visual evidence, we evaluate generated projects along four dimensions---Executability, Specification Realization, Code Quality, and User-Facing Quality. Across 14 LLMs and 3,330 generated projects, we find that high executability does not imply faithful specification realization. Component-level analysis further shows that models perform substantially better on Game Element Modeling than on Rule and Mechanism Modeling or Goal and Termination Modeling, indicating that faithfully implementing game rules and termination logic remains a major challenge.

Comment: 37 pages, 11 figures

arXiv abs page · PDF · same-day batch