PaperScope
LIVE · 2026-09-03 05:40 UTC

The Privacy-Hallucination Tradeoff in Differentially Private Language Models

Krithika Ramesh, Krishna Pillutla, Danish Pruthi, Anjalie Field

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.00492 v1
Category
Submitted
2026-08-31

Abstract

Both privacy and factual accuracy are paramount in high-stakes domains like healthcare. Concerningly, we uncover and investigate a privacy-hallucination tradeoff in differentially private (DP) language models. First, we empirically show that models pre-trained or fine-tuned with DP tend to produce more hallucinations than non-DP counterparts, with increased severity as the privacy budget grows stricter. Second, we investigate model properties driving this tradeoff, demonstrating that DP mechanisms flatten output distributions, potentially redistributing probability mass toward factually incorrect alternatives. Third, through experiments where we control fact frequency in training data, we characterize how information frequency can reduce hallucination risks in DP models. Overall, our findings underscore the need for more nuanced privacy-preserving interventions that offer rigorous privacy guarantees without compromising factual accuracy.

Comment: Accepted to EMNLP 2026 (Findings)

arXiv abs page · PDF · same-day batch