PaperScope
LIVE · 2026-09-15 05:40 UTC

CiteGuard-RAG: A Validation-Centered AI System for Evidence-Grounded Question Answering

Sumit Barua, Guan Hong, Halil Dursunoglu, Charles Rodgers, Alvis Fong

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.15830 v1
Category
Submitted
2026-09-14

Abstract

Retrieval-augmented generation (RAG) can improve access to complex information; however, retrieving evidence alone does not ensure that answers are grounded, citation-valid, or appropriately refused. This paper introduces CiteGuard-RAG, a validation-centered AI system for evidence-grounded question answering. The system integrates hybrid semantic-lexical retrieval, citation-constrained generation, sentence-level grounding validation, and single-pass regeneration. Validation is used at runtime to determine whether a candidate answer should be accepted, refused, or regenerated before final delivery. CiteGuard-RAG is evaluated on 400 questions across a controlled housing-law dataset, PrivacyQA, and CUAD. In the controlled evaluation, it achieves 99.1% retrieval accuracy, 98.3% grounded-answer accuracy, and 98.3% citation validity, with no validation-detected hallucinations. Ablation results show that grounded-answer accuracy drops sharply when validation is removed, even when retrieval accuracy remains unchanged. External evaluation shows that while citation validity remains strong, evidence utilization, span alignment, and refusal calibration become harder under domain shift. These findings indicate that trustworthy RAG systems require explicit validation between retrieval and final answer delivery. CiteGuard-RAG provides a practical architecture for linking retrieval, generation, citation checking, abstention, and regeneration in high-stakes information access.

Comment: Submitted to Engineering Reports. 22 pages, 2 figures, 12 tables

arXiv abs page · PDF · same-day batch