Representation Risk in Pretrained Image Encoders
Ardyn Nordstrom, Morgan Nordstrom, Vamuyan Sesay, Matthew D. Webb
Abstract
Applied researchers increasingly convert images into features with pretrained encoders, then use those features in a downstream prediction model. The encoder is often treated as an implementation detail. We show that it can instead be a consequential source of model uncertainty. We call this uncertainty representation risk: plausible pretrained encoders map the same images into different feature spaces and can yield sharply different out-of-sample conclusions from predictive performance. We compare ten modern and legacy frozen encoders across applications involving house prices, racehorse performance, breast-cancer histology, chest radiographs, continuous facial age, and rice disease. With common dimension control, heads, and group-safe splits, validation selects SigLIP 2 for houses, raising test $R^2$ from 0.396 for ResNet50 to 0.629, and DINOv2 for horses, raising $R^2$ from 0.029 to 0.105. No encoder is best in every task. Candidate procedures are constructed using training data and compared on a separate validation partition. The selected procedure reaches 0.658 for houses and 0.979 accuracy for pneumonia. Fixed-split gains are small for horses and rice, while repeated partitions reveal instability in horse feature union. Continuous age selects SigLIP 2 at 4.786 years MAE. The principal representation gaps persist with neural heads, similarly sized DINOv2 and ViT models, and limited adaptation. These results support a simple workflow: benchmark plausible representations, select on locked validation data, combine only when separate validation evidence justifies the additional cost, and report paired and split-level uncertainty. We implement this workflow in LOOKAGAIN-ML, the software package used to conduct the analyses in this paper.