PaperScope
LIVE · 2026-09-29 05:40 UTC

From PDF to Evidence: Structure-Aware Retrieval for Clinical Practice Guidelines

Xingyu Lin, Dehui Du

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.33447 v1
Category
Submitted
2026-09-27

Abstract

Guideline documents are published as unstructured PDFs whose evidence is locked in visual structures---tables, flowcharts, and graded recommendations---that standard retrieval pipelines flatten into fixed-size text chunks. We cast evidence access as a document image analysis problem: parse each page image into typed structural elements, then retrieve structure-aware evidence units that follow the document's own layout (sections, table rows, flowchart paths, graded recommendations), each keeping its structural context so a result points to a specific element rather than a page. On 26 clinical practice guidelines from 9 sources (3,619 pages, Chinese and English) with 199 evidence queries, structure-aware units rank the gold element first under BM25, dense, and hybrid retrieval (hybrid Element Hit@1 of 0.382), with a significant element-level ranking gain over per-element OCR text (MRR_e +0.107, p=0.002; the Hit@5 gain is directional, p=0.17), while matching page-level recall (Page Hit@5 0.879 vs. 0.889, p=0.75) at 3.8x less context and clearly outperforming a ColPali visual-RAG baseline (PH@5 0.497).

Comment: 5 pages, 2 figures, 5 tables. Submitted to ICASSP 2027

arXiv abs page · PDF · same-day batch