PaperScope
LIVE · 2026-09-29 05:40 UTC

Scalable In-Domain Self-Supervised Foundation Model for Dense Representation Transfer in High-Resolution Plant Imaging

Junlin Guo, Sharmin Majumder, Isaac Lyngaas, John Lagergren, Xiao Wang

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.32183 v1
Category
Submitted
2026-09-26

Abstract

High-resolution plant imaging enables detailed characterization of plant morphology, but dense scientific analysis remains limited by costly pixel-level annotations, large image pixel dimensions, and substantial variation in imaging conditions. This work proposes a scalable in-domain self-supervised pretrained foundation model for high-resolution, high-pixel-dimension multi-species plant imagery. A masked autoencoder with a ViT backbone is pretrained on more than 10 million multi-view plant image tiles using distributed training. Following scalable pretraining, the learned foundation-model representations are comprehensively benchmarked across fine-grained dense prediction and coarse global feature recognition, with particular emphasis on limited supervision and realistic downstream imaging conditions. This work focuses on the domain gap of existing foundation models in dense feature representation and transfer. Through extensive experiments involving limited annotations, cross-view variation, and resolution degradation, the in-domain FM achieves a Mean Dice of 0.8686 and a Pooled Dice of 0.8959, outperforming an MAE counterpart pretrained on large-scale natural-image data by 0.0694 and 0.0613, respectively. The results further indicate that increasing pretraining scale produces consistent improvements in dense feature transfer. Overall, these findings suggest that scaling in-domain self-supervised pretraining can reduce the domain gap and improve transferable dense representations for high-pixel-dimension scientific imaging.

arXiv abs page · PDF · same-day batch