PaperScope
LIVE · 2026-10-06 05:40 UTC

VIGIL: Verifier-Informed Gated Improvement Loop for Spreadsheet Question Answering

Kang Li, Lu He, Sandarsita Guntupalli

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.04287 v1
Category
Submitted
2026-10-03

Abstract

Enterprise agents should improve from delayed feedback without allowing every correction to rewrite system behavior. We study continual harness learning for corpus-level spreadsheet question answering. Building on FiCo (Find-then-Compute), a static retrieval-and-execution backbone, we introduce VIGIL (Verifier-Informed Gated Improvement Loop). Within a question, VIGIL verifies and repairs diversely prompted Structured Query Language (SQL) candidates. Across episodes, delayed labels update only a contract calibrator and query/result-column selector. The base model, prompts, retriever, and recorded candidate pool remain fixed in the continual-learning protocol. With the gold workbook and expected type supplied, the ungated and dual-gated full-replay variants reach 82.4% and 82.0% forward accuracy over 79 documents, from a 75.8% static baseline. The dual gate has the larger retrospective gain (2.7 versus 2.3 points), and its 6.3-point forward gain has a 95% t-interval of 5.6-6.9. In a separate stricter split that excludes 16 documents and 380 questions from fitting and online promotion, the accuracy-only gate raises mean held-out accuracy across ten final harnesses from 80.3% to 86.7%. Calibrator-only adaptation gains 5.9 points, close to the combined 6.4-point gain. Yet three of 26 gate-approved updates reduce held-out accuracy relative to their incumbents, so replay-buffer non-regression does not imply held-out non-regression. MiMoTable and external-task case studies test within-task verification and reuse the same promote-or-retain discipline. Overall, the results support bounded, auditable harness adaptation while revealing where finite replay gates fail to generalize beyond their promotion buffers.

arXiv abs page · PDF · same-day batch