M3-Score: Fidelity, Memorization and Coverage as Separate Axes for Evaluating Generative Radiology Image Models
Sathiyamohan Nishankar, Pubudu Sanjeewani, Asanka Perera
Abstract
Quantitative evaluation of generative models for radiology remains challenging. Clinically relevant structures are often small and infrequent, feature spaces learned from natural images may represent them poorly, and a single summary score cannot distinguish limited fidelity from limited diversity. This study proposes the Medical Multi-axis Maximum Mean Discrepancy score (M3-Score), an evaluation framework based on RadioDINO-s16, a frozen vision transformer pretrained on radiology images. M3-Score reports three complementary axes computed at pre-specified encoder depths: \emph{fidelity}, measured by an unbiased multi-bandwidth radial basis function (RBF) MMD$^2$ at the final block; \emph{memorization}, measured by nearest-neighbor distances at 75\% depth; and \emph{coverage}, defined as the fraction of real images with a generated neighbor within their $k$-nearest-neighbor radius at 33\% depth. Reference sets are sampled across subjects to limit the influence of correlated slices. On BraTS brain MRI, the fidelity axis ordered five comparison sets of increasing severity (Spearman $ρ= 1.00$), and a subject-disjoint real set yielded $\mathrm{MMD}^2 = 0$ (permutation $p = 1$). An unconditional denoising diffusion probabilistic model achieved $\mathrm{MMD}^2 = 0.073$ (95\% confidence interval $[0.071, 0.080]$) but covered only 38\% of the real distribution. Under progressive mode dropping, $k$-NN manifold recall increased at all twelve encoder blocks, whereas the proposed coverage estimator decreased monotonically ($ρ= -1.00$). RadioDINO-s16 features separated real brain MRI from generated samples with a ROC-AUC of 0.819, compared with 0.555 for InceptionV3 and 0.582 for CLIP. Across a twentyfold range of sample sizes, the mean M3 value varied by a factor of 1.05, compared with 2.52 for the Fréchet Inception Distance.