PaperScope
LIVE · 2026-09-03 05:40 UTC

On the Recoverability of Private Information Unlearning in Large Language Models

Shicheng Hu, Runzhi Tian, Ziqiao Wang, Yongyi Mao

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2608.29943 v1
Category
Submitted
2026-08-30

Abstract

Large language models (LLMs) can memorize sensitive information, raising serious privacy concerns. Machine unlearning offers a potential solution to remove such information, but it remains unclear whether existing methods truly erase it or merely hide it within the model. A key challenge is quantifying the persistence of sensitive data under a unified evaluation framework. To address this, we construct a synthetic dataset containing fake private information and propose a white-box auditing framework to systematically assess whether claimed-forgotten information is genuinely removed. Using this framework, we evaluate five existing unlearning methods and find that a simple "inverse greedy" decoding -- selecting the least likely token at each step -- can recover supposedly forgotten private information. Our results reveal that current unlearning approaches often fail to fully eliminate sensitive information, highlighting the need for more reliable methods to ensure privacy in deployed LLMs.

arXiv abs page · PDF · same-day batch