PaperScope
LIVE · 2026-09-29 05:40 UTC

Is H&E Image-to-Spatial Transcriptomics Simpler Than It Looks?

Duc T. Nguyen, Thanh Ha Do, Phuong M. Cao, Hieu Pham

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.32857 v1
Category
Submitted
2026-09-26

Abstract

Predicting spatial gene expression from routine H&E histology offers a scalable route toward spatial molecular profiling. Recent work has pursued increasingly sophisticated architectures to capture spatial context and richer expression structure. At the same time, simple estimators have shown strong performance in several studies, but what they already solve and where additional complexity is needed remain unclear. We study this behavior through the structure of prediction error under the mean-squared error (MSE) objective. Differences in average expression across genes can account for a substantial part of aggregate prediction performance, while a key unresolved error lies in recovering variation within each slide. Decomposing MSE into slide-level and within-slide components, we find that the within-slide component has lower residual-normalized parameter sensitivity in controlled neural experiments. This motivates Component-Guided Loss (CGL), which increases supervision of the within-slide component. CGL-Linear is a closed-form affine instantiation that achieves overall state-of-the-art performance across HEST-1k cohorts and gene-panel sizes. The same within-slide supervision improves existing neural models. These results suggest that substantial gains can come from aligning the training objective with prediction-error structure rather than increasing model complexity.

arXiv abs page · PDF · same-day batch