PaperScope
LIVE · 2026-09-09 05:40 UTC

Are Image Generators Zero-Shot Perceivers? A Rigorous Evaluation

Shangzhe Di, Zhaokai Wang, Weidi Xie

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.07884 v1
Category
Submitted
2026-09-07

Abstract

Recent work, such as Vision Banana, shows that lightweight instruction tuning can enable an image generator to achieve state-of-the-art performance across multiple visual perception tasks. Motivated by this perspective, we ask how far image generators can go on public visual perception benchmarks in a zero-shot setting. We introduce ProbeGen, a benchmark for zero-shot generative perception that casts monocular depth estimation, referring/reasoning segmentation, and object counting as conditional generation tasks specified through text prompts, and compares 20 models in total---including proprietary and open-weight image generators, specialist perception models, and MLLMs---across 11 published benchmarks. We observe that pretrained image generators show measurable zero-shot perceptual competence, but with a clear trade-off: specialist models remain stronger for in-distribution accuracy and efficiency, while generative models are often more robust under distribution shift and better at compositional semantic reasoning. We hope this study helps establish zero-shot generative perception as a meaningful research direction and provides a useful foundation for future work at the intersection of visual generation and understanding.

Comment: Accepted to BMVC 2026

arXiv abs page · PDF · same-day batch