PaperScope
LIVE · 2026-09-03 05:40 UTC

Training seeds and model-selection stability in recommender-system evaluation

Juan Manuel Rodriguez, Oleg Lesota, Antonela Tommasel

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.02499 v1
Category
Submitted
2026-09-02

Abstract

Recommender-system experiments often rely on a single random training seed, assuming that run-to-run stochasticity has limited impact on evaluation conclusions. This assumption is risky, as a training seed may influence several algorithm-dependent mechanisms, including parameter initialization, mini-batch ordering, dropout, masking, latent sampling, and training-time negative sampling. We examine this assumption by fixing the data partition and varying the training seed across hyperparameter configurations. We analyze seed effects at three levels: user-level metric sensitivity, validation-based model selection and recommendation-list agreement. Results show that seed variation is often detectable. Its impact depends on whether configurations are clearly separated, whether validation results transfer to test, and whether similar scores lead to similar top-$k$ lists. Findings suggest that reporting single-seed results can overstate the stability of recommender system evaluation, and that training seeds should be treated as part of the evaluation protocol rather than as incidental implementation noise.

Comment: Accepted RecSys 2026 (Research&Practice Notes)

arXiv abs page · PDF · same-day batch