PaperScope
LIVE · 2026-10-01 05:40 UTC

Evidence First, Arithmetic Second: A System Report and Failure Analysis for DocSem

Divya Godara, Sachin Gupta

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.39013 v1
Category
Submitted
2026-09-30

Abstract

EVICALC, our system for the DocSem shared task, achieved 8.61% joint accuracy on 1,730 tasks in the official final test evaluation. It reads a PDF, selects a passage, asks a language model to write an arithmetic expression, and evaluates that expression in local code. Saved intermediate results support inspection of failures. A separate public-validation run achieved 92.17% answer accuracy and 1.00 evidence F1. The configurations and metrics differ, so these scores are not a controlled comparison. Our manual, post-hoc analysis is descriptive: in one inspected case, optical character recognition (OCR) and block grouping merged the relevant passage into another block, and the system answered from unrelated text. An exploratory study of reading page images on 100 documents returned evidence identifiers for only 22 documents. These descriptive findings motivate further evaluation; they do not establish the causes of the overall score.

Comment: 5 pages, 1 figure, 2 tables. Accepted as a shared-task system paper at DocInsights 2026, co-located with EMNLP 2026

arXiv abs page · PDF · same-day batch