PaperScope
LIVE · 2026-10-08 05:40 UTC

OverLay++: Dense-Overlap Layout-to-Image Generation Dataset

Shivansh Aggarwal, Shresth Grover, Divyansh Srivastava, Haiyang Xu, Bingnan Li, Xiang Zhang, Ethan J. Armand, Chuan Li, Jianwen Xie, Zhuowen Tu

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.09071 v1
Category
Submitted
2026-10-06

Abstract

Layout-to-Image generation has made substantial progress in spatial and object-level control. However, existing methods still struggle with complex scenes containing many overlapping and interacting objects. We argue that training data is a particular bottleneck: existing datasets lack examples with dense, complex object interactions. To address this gap, we introduce OverLay++, a large-scale Layout-to-Image dataset with structurally complex scenes. OverLay++ contains approximately 500K images with an average of 6.6 objects per image, exceeding existing datasets by 1.67 times in annotation density. Beyond annotation density, OverLay++ provides rich semantic detail with object captions over six times longer than in current datasets. Our dataset generation pipeline is simple and produces dense, overlapping object annotations with rich per-object captions. Across multiple benchmarks, state-of-the-art Layout-to-Image methods trained on the OverLay++ dataset show consistent improvement and faster convergence, demonstrating the importance of dense, overlap-aware, and caption-rich supervision for controllable image generation.

Comment: Accepted at NeurIPS 2026, Evaluations & Datasets Track. Project website: https://mlpc-ucsd.github.io/OverLayPP . Dataset: https://huggingface.co/datasets/mlpcucsd/OverLayPP

arXiv abs page · PDF · same-day batch