PaperScope
LIVE · 2026-10-07 05:40 UTC

Can phenotypic activity be predicted without experimental readouts?

Télio Cropsal, Rocío Mercado

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.07997 v1
Category
Submitted
2026-10-06

Abstract

Molecular encoders contrastively pretrained on paired molecule-morphology data, such as CLOOME and CellCLIP, have been proposed as cheap surrogates for phenotypic prediction, avoiding the need to run a Cell Painting assay. We evaluate this idea for these molecular encoders under a protocol designed to control for two confounds that can inflate apparent performance: leakage across an encoder's own pretraining boundary, and the correlation between phenotypic activity and cytotoxicity. Testing six representations, including a non-pretrained MLP control matching CLOOME's input and layer count, on two distinct Cell Painting screens, we find that once these confounds are controlled for, the pretrained molecular encoders show no clear advantage over plain physicochemical descriptors, and that toxicity is generally easier to predict than phenotypic activity across representations. Our results suggest leakage-aware, confound-controlled evaluation should be standard practice before phenotype-pretrained encoders are trusted as surrogates for phenotypic drug discovery.

Comment: Accepted to the ML4Molecules: Agentic Systems for Molecular Sciences Workshop at the 40th Conference on Neural Information Processing Systems (NeurIPS 2026)

arXiv abs page · PDF · same-day batch