PaperScope
LIVE · 2026-09-30 05:40 UTC

Look Closer: Patch-wise Supervision for AI-Generated Image Detection

Zhida Zhang, Tao Wu, Siyu Liu, Jie Cao

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.37937 v1
Category
Submitted
2026-09-29

Abstract

How much of an image does a detector need to see? Small RGB regions can retain useful evidence of image synthesis even when they reveal little of the full scene. Motivated by single-patch detection, we study patch-wise supervision: a shared backbone classifies explicit crops, each crop receives its own loss, and patch probabilities are averaged only at inference. The procedure requires neither handcrafted residual filtering nor a learned image-level fusion module. Experiments span single-patch selection, multiple generator collections, and four CNN and Transformer backbones. On GenImage, the reported patch-wise variants improve average accuracy over their whole-image counterparts across all four backbones. Comparisons of supervision granularity, source resolution, crop size, and inference coverage further characterize the approach, while post-processing tests and difficult-image evaluation reveal its limitations. The historical experiments include evaluation-based model selection, so their scores are not presented as a uniformly selected leaderboard comparison. Overall, the study identifies explicit local input and patch-level supervision as a simple, useful combination for investigating generalizable AI-generated image detection.

Comment: 29 pages, 11 figures, 28 tables. Code: https://github.com/LF-Jade/look-closer

arXiv abs page · PDF · same-day batch