PaperScope
LIVE · 2026-09-09 05:40 UTC

CIVI: A Framework for Diagnosing Search Agent Failures in Civic Information

Dingying Liu, Yunshun Zhong, Wentao Zhang, Yiyuan Li

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.08094 v1
Category
Submitted
2026-09-08

Abstract

Large Language Models are increasingly deployed in public-sector settings, where incorrect guidance can cause irreversible harm. We introduce CIVI, the first framework for diagnosing search agent failures in civic information. Its benchmark instantiation jointly spans cross-national, interjurisdictional government contexts (federal, state, and local) and functional categories from an internationally adopted United Nations standard. We evaluate ten frontier search agents and find that none matches an attentive human baseline. Alongside accuracy, CIVI measures search invocation rate, selective no-search accuracy, and how often agents cite authoritative government sources. To perform this diagnosis, we introduce ARISE, which decomposes agentic search failures into four mutually exclusive modes, isolated via source-injection ablation. ARISE attributes 72.1% of all observed failures to retrieval-bound causes rather than to gaps in the models' parametric knowledge.

arXiv abs page · PDF · same-day batch