PaperScope
LIVE · 2026-09-29 05:40 UTC

VD-DeepStack: Bridging Visual Comparison and Language Reasoning for Few-Shot Anomaly Detection

Mengyang Zhao, Zhuolin He, Haiyang Yu, Yuxuan Liang, Yifang Xu, Yuchuan Wu, Xiaolei Chen, Zhengtao Yao, Fan Shi, Yang Liu, Bin Li, Xiangyang Xue

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.34949 v1
Category
Submitted
2026-09-28

Abstract

Few-shot visual anomaly detection is fundamentally a visual comparison task, requiring fine-grained inspection of a query against normal references. Many recent methods based on large vision-language models (LVLMs) emphasize comparative reasoning through language chain-of-thought. Yet discrete, abstract descriptions may underrepresent dense, fine-grained visual differences, leaving a gap between visual comparison and its expression in language. To address this gap, we propose Visual Difference DeepStack (VD-DeepStack), which explicitly conditions language reasoning on query-reference visual differences. Specifically, we fuse DINO features with the LVLM visual hierarchy to strengthen fine-grained representations, then construct dense difference evidence from residuals between query features and softly matched reference features. The difference-evidence path injects spatially weighted difference vectors into query-image states at multiple decoder depths, while an auxiliary visual-context path provides fine-grained appearance information to support their interpretation. Experiments on 4 industrial and 2 medical anomaly benchmarks demonstrate substantial improvements in few-shot anomaly detection over baselines relying on textual comparative reasoning. These results support mitigating the visual comparison-reasoning gap through the joint design of comparison representations and their integration into the decoder. Code will be released upon acceptance.

arXiv abs page · PDF · same-day batch