PaperScope
LIVE · 2026-09-29 05:40 UTC

AmbiModBench: Benchmarking Gene Perturbation Prediction Beyond Shared Responses

Sikai Huang, Zhiwen Yang, Kai Yu, Jiayuan Chen, Stan Z. Li

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.32527 v1
Submitted
2026-09-26

Abstract

Predicting cellular responses to genetic perturbations helps prioritize experiments in single-cell genomics, where exhaustive measurement is infeasible. While computational models increasingly predict these responses, three evaluation deficiencies obscure what their scores demonstrate. First, absolute metrics cannot separate target-specific predictions from a shared background response. Second, common metrics remain high under gene shuffling, so gene-level accuracy is never verified. Third, a score at one training size says nothing about coverage, which depends on representation-space proximity and response-constraining power. We propose AmbiModBench, a specificity-aware, gene-resolved and coverage-aware benchmark. It pairs every score with a training-mean reference fitted on the same split, screens each readout by gene-coordinate permutation, and links embedding distance to response variation. Across K562, RPE1 and Norman, strong absolute scores largely reflect shared background rather than target-specific learning. Widely used readouts track response magnitude distributions rather than the affected genes. Detectable gain follows representation-space coverage rather than training-set size. Nonetheless, on RPE1 the protocol yields a reproducible target-specific gain across five additional splits and three gene selections, which absolute scores alone cannot distinguish from shared background.

Comment: 21 pages, 4figures

arXiv abs page · PDF · same-day batch