PaperScope
LIVE · 2026-09-29 05:40 UTC

InterTab: Interleaved Visual-Structure Alignment for Multi-Modal Table Reasoning

Hanqian Li, Sirui Huang, Chen Ling, Jungang Li, Yu Huang, Kening Zheng, Yonghua Hei, Xiangrong He, Shiyi Wang, Pengcheng Zhu, Dongnan Liu, Wei Zhou, Linjian Mo, Nai Ding, Xuming Hu

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.32660 v1
Category
Submitted
2026-09-26

Abstract

Table images preserve structural information that are often lost in text serialization, and reasoning over them requires locating relevant rows, columns, and cells step by step. Current multimodal large language models (MLLMs) encode the whole image once before reasoning, so they cannot pick up row-, column-, and cell-level evidence as the question unfolds. Encoder-side table structure and generic interleaved visual chain-of-thought still do not bind each reasoning step to that structure. We propose \textbf{InterTab}, an \textbf{Inter}leaved structure-aware framework for CoT reasoning over \textbf{Tab}le images, interleaves chain-of-thought with tool calls that crop structure-aligned table regions. First, we build InterTab-22K, includes reasoning trajectories in which each step is tied to both a structural location and a bounding box. InterTab is trained in two stages: supervised structure-aware alignment (SSA) on InterTab-22K teaches the model to interleave reasoning with structure-aligned crops, and active localization optimization (ALO) further optimizes answer correctness, localization IoU, and output format, while penalizing missing or excessive tool calls. Experiments on nine table benchmarks show that InterTab improves the average accuracy of its backbone from 68.28% to 73.17% and achieves the best average performance among all compared methods. Code and data will be released soon.

arXiv abs page · PDF · same-day batch