PaperScope
LIVE · 2026-10-08 05:40 UTC

SpatialUQ: Post-Hoc Uncertainty Quantification from Spatial Consistency in Black-Box Vision Models

Md Kawsher Mahbub, Milon Biswas, Mirza Niaz Morshed, Wei Yu

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.09498 v1
Category
Submitted
2026-10-07

Abstract

Clinical vision models are often deployed as frozen black boxes with no access to internals, retraining, or ground truth at inference time. We introduce \textbf{SpatialUQ}, a post-hoc uncertainty method using only output probabilities. It measures the Jensen-Shannon divergence between the global prediction and the mean of five fixed spatial crops in six deterministic forward passes. The premise is simple, trustworthy predictions are spatially consistent. On NIH ChestX-ray14 (DenseNet-121, $N{=}25{,}596$), our Multicrop Uncertainty Score (MUS) reaches $0.784$ failure-detection AUC versus $0.664$ for MC-Dropout ($p{<}10^{-6}$) at one-fifth the compute, with native calibration ($\text{SCE}{=}0.049$ vs.\ $0.127$ for $\ell_1$), the best-calibrated among methods above 0.78 AUC. A supervised fusion of MUS with entropy, confidence, and $\ell_1$ reaches $0.832$, outperforming a five-member ensemble ($0.813$). MUS scales with model quality, reaching $0.899$ with BiomedCLIP ($ρ= 0.846$), while this relationship remains meaningful in-distribution ($ρ= 0.523$) but breaks down under severe distribution shift (VinBigData, $ρ= 0.027$). MUS is well-suited to diffuse findings but is less dependable for small focal lesions such as nodules. Code and experimental materials are publicly available at https://huggingface.co/datasets/kawsher11/SpatialUQ.

arXiv abs page · PDF · same-day batch