PaperScope
LIVE · 2026-10-01 05:40 UTC

Consensus-Aware Multi-Source Fusion for Reference-Guided Camouflaged Object Detection

Junyang Xia, Luocheng Zhang, Wenwen Pan, Chifeng Zhu, Yang Yang, Xinchun Liu, Jiajun Ding

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.38747 v1
Category
Submitted
2026-09-30

Abstract

Reference-guided camouflaged object detection aims to segment a target whose visual appearance closely resembles its surroundings by exploiting auxiliary reference samples. The task remains difficult because reference samples contain inconsistent target cues, while generic visual representations are not inherently aligned with the target specified by the references. To handle these problems, we present a consensus-aware multi-source fusion framework. Reference-Conditioned Dual-Backbone Fusion (RCDF) couples trainable PVTv2 query features with frozen DINOv3 representations and uses reference-conditioned correlation to select foundation-model evidence before multi-scale fusion. The framework also aggregates multiple references through cross-reference consensus aggregation and injects reference information at semantic depths matched to the query features. Extensive experiments demonstrate the effectiveness of the proposed method. The results further show that reference consensus, target-conditioned foundation features, and hierarchical decoding provide complementary improvements under the evaluation protocol. The source code will be made publicly available upon acceptance.

Comment: 20 pages, including 2 pages of appendix; 9 figures and 4 tables

arXiv abs page · PDF · same-day batch