PaperScope
LIVE · 2026-10-07 05:40 UTC

RAG-PIBench: A Leakage-Aware Benchmark for Prompt-Injection Detection in Trustworthy RAG Systems

Niveen O. Jaffal, Ahmet Yuksel, David Mohaisen

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.08571 v1
Category
Submitted
2026-10-06

Abstract

Retrieval-Augmented Generation (RAG) systems are vulnerable to prompt-injection attacks embedded in retrieved content. We introduce RAG-PIBench, a benchmark for RAG-style prompt-injection detection containing 4,876 contextual examples across frozen train, validation, and protected-test splits. Using a leakage-aware construction pipeline and strict evaluation protocol, we compare keyword-based, semantic-reference, TF-IDF, and transformer-based detectors. DistilBERT achieves the best protected-test performance (F1 = 0.896, PR-AUC = 0.968), while TF-IDF SVM and logistic regression remain competitive. Our results demonstrate the value of leakage-aware benchmark design and strong sparse baselines for reliable prompt-injection detection in RAG systems.

Comment: 19 pages, 3 figures, 8 tables

arXiv abs page · PDF · same-day batch