PaperScope
LIVE · 2026-10-05 05:40 UTC

Benchmarking Literature Retrieval for a Model Organism: A Dictyostelium Case Study

Yun Wang, Gad Shaulsky, Tomaž Curk, Blaž Zupan

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.03130 v1
Category
Submitted
2026-10-02

Abstract

Biological literature retrieval systems are often developed and evaluated using broad biomedical corpora and general-purpose search tasks. However, many curated knowledge bases operate in narrower model-organism domains, where the literature is sparse and terminology is organism-specific. We introduce a retrieval benchmark from dictyBase for Dictyostelium, a model organism in cell and developmental biology. The benchmark consists of curator-generated biological queries linked to PubMed-indexed articles, together with structured gene annotations. Using this benchmark, we study three factors in niche biological retrieval: cross-encoder reranking, gene-aware query expansion, and abstract-only versus full-text retrieval. We report that reranking and gene-aware query expansion improve retrieval selectively: reranking is most useful when the model is well suited to biological evidence matching, whereas curated annotations help clarify compact biological queries by reducing vocabulary mismatch. Full-text chunks substantially improve retrieval when abstracts omit supporting evidence, increasing both candidate recall and top-rank performance, although these cases are harder than queries supported by abstracts. Data and code are publicly available at https://github.com/fulaibaowang/dictycite, and the benchmark dataset is additionally archived on Zenodo.

Comment: 15 pages, 5 figures. Submitted version (before peer review) of a paper accepted at Discovery Science 2026 (DS 2026); to appear in the Springer proceedings. Code and data: https://github.com/fulaibaowang/dictycite ; dataset: https://doi.org/10.5281/zenodo.20308282

arXiv abs page · PDF · same-day batch