Rank-Consistent Set Reasoning for Co-Salient Object Detection
Yuan Xiang, Matteo Rossi, Yingzhou Chen
Abstract
Co-salient object detection (Co-SOD) requires a model to find foreground regions that are salient in individual images and supported by the image group. We present \emph{Rank-Consistent Set Reasoning} (RCSR), a supervised dense-prediction framework that models a group as an unordered set rather than as a sequence of images or a semantic label. The core idea is to rank how strongly each spatial region agrees with a small collection of learned group slots at every image scale, and to aggregate these ranks with a robust trimmed statistic. This suppresses accidental pairwise matches and prevents one atypical group member from dominating the shared representation. A set encoder builds group slots directly from multi-scale visual features, while a rank-consistency gate measures whether the ordering of candidate regions is stable across group members. The gated slots are decoded jointly with per-image features to produce co-saliency maps. The model contains no natural-language branch, no open-vocabulary detector, and no external segmentation model. We further introduce a group permutation objective and hard-distractor augmentation so that the model learns the properties of a set-level target rather than memorizing image order or isolated visual saliency. We formulate an evaluation protocol for CoCA, CoSal2015, and CoSOD3k, together with tests of group-size robustness, distractor rejection, order invariance, and cross-dataset transfer.