PaperScope
LIVE · 2026-09-29 05:40 UTC

BudgetVerify: Budget-Tiered Verification for Financial QA

Janet Jenq, Hongda Shen

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.33052 v1
Category
Submitted
2026-09-27

Abstract

Financial question answering often requires precise numerical extraction, unit handling, and arithmetic over tables and text, but applying expensive verification uniformly wastes test-time compute. We propose BudgetVerify, a budget-tiered generator-verifier framework that routes each generated answer to one of three verification tiers: no verification, lightweight check-and-revise, or higher-cost solve-first-then-compare verification. The router is trained from offline correctness and token-cost outcomes and, at test time, selects a verification tier using information available before verification, including the question, context statistics, the generated answer, and associated generator metadata. The selected tier either returns the generated answer directly or invokes the corresponding verifier. Across six commercial and open-weight base models, BudgetVerify consistently produces more efficient accuracy-cost Pareto frontiers than fixed verification policies by selectively allocating stronger verification only when it is useful. Although absolute performance varies across models, these efficiency gains and the resulting qualitative frontier shape are consistent across generator models.

arXiv abs page · PDF · same-day batch