PaperScope
LIVE · 2026-09-03 05:40 UTC

Text-Driven Artistic Staging: Pose, Lighting, and Camera References from Paintings

Yunge Wen

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2608.28823 v1
Category
Submitted
2026-08-28

Abstract

Artists coordinate human pose, illumination, and camera placement to convey narrative and emotion, but existing generative methods typically model these elements independently. We introduce text-to-editable 3D staging, a task that jointly generates human poses, a dominant light, and a camera configuration from an affective description. We construct 11,911 text--staging pairs from 2,328 figurative paintings by reconstructing SMPL bodies, estimating low-frequency illumination, recovering camera parameters, and pairing each scene with ArtEmis descriptions. We train a flow-matching transformer that supports variable numbers of figures and produces multiple staging alternatives for each prompt. On held-out descriptions, the model achieves 32.2\% retrieval R@1, compared with 16.6\% for CLIP-based nearest-neighbor retrieval, while approximately preserving corpus-level diversity. These results demonstrate the feasibility of generating editable, emotionally conditioned 3D staging references from text.

arXiv abs page · PDF · same-day batch