PaperScope
LIVE · 2026-10-06 05:40 UTC

EnvDreamer: Large-Scale Multimodal-to-Environment Generation for Embodied AI

Kabir Swain, Sijie Han, Antonio Torralba

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.04301 v1
Category
Submitted
2026-10-03

Abstract

Large datasets and high capacity models have accelerated progress in vision and language. This work introduces a platform aimed at bringing comparable gains to embodied learning, world models, and robotics. We present EnvDreamer, a framework that uses large language and vision language models to generate Unreal Engine 5 environments for embodied AI and robot training. EnvDreamer enables sampling of large, diverse, interactive, customizable, and validator passed virtual environments for training and evaluation across navigation, interaction, and manipulation. We illustrate the platform with a large set of generated scenes and simple baselines. Policies trained on EnvDreamer generated environments, without explicit mapping or human task supervision, achieve competitive results on multiple embodied benchmarks spanning navigation, rearrangement, and manipulation. EnvDreamer also supports image-conditioned reconstruction for real-to-sim studies. Finally, we release EnvDreamer-20k, a dataset of 20,000 validator passed environments with task programs, scene graphs, trajectories, and metadata to support reproducible benchmarking.

arXiv abs page · PDF · same-day batch