PaperScope
LIVE · 2026-09-21 05:40 UTC

PrismAlign: Prior-Steered Multi-View VLM Alignment for Hallucination-Robust Table OCR

Guangyi Liu, Qianjun Huang, Boyu Hou

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.21351 v1
Category
Submitted
2026-09-18

Abstract

Table extraction suffers from frequent structural errors and semantic hallucinations. We propose PrismAlign, a multi-VLM framework aligning diverse visual perspectives to resolve ambiguity. It integrates priors of table logic to assess output plausibility, decoupling structural alignment from cell content alignment. A Bayesian decision strategy maximizes alignment accuracy by exploiting the correlation between extraction errors and computable rule violations. Evaluated on open-source and custom VLMs, PrismAlign reduces hallucinations and achieves state-of-the-art performance on OmniDocBench 1.5, as well as on the table category of CC-OCR and PureDocBench.

Comment: Accepted by EMNLP industry track

arXiv abs page · PDF · same-day batch