PaperScope
LIVE · 2026-09-03 05:40 UTC

Position: Privacy Is a Claim, Not a Property of Synthetic Data

Jiachen Zhao, Antonia Januszewicz, Taeho Jung

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.01273 v1
Category
Submitted
2026-09-01

Abstract

Synthetic data has become a common component of machine learning research. While widely adopted, its use in privacy-sensitive contexts has quietly shifted from a claim of residual inference risk under stated assumptions to an appearance-based property inferred from data generation itself. In this position paper, we argue that this shift reflects an implicit change in community standards for what counts as sufficient privacy evidence, rather than a misunderstanding of well-established privacy principles. Drawing on an empirical analysis of recent publications across major ML venues, we show that synthetic data is frequently used in privacy-sensitive settings without explicit articulation of threat models, inference risks, or falsifiable privacy claims. As a result, privacy assurance often remains implicit, difficult to verify, and unevenly distributed, with heightened exposure for rare and minority records. We argue for treating privacy as an explicit, evidence-based scientific claim and recommend that ML venues adopt norms requiring privacy-relevant assertions to be clearly scoped, testable, and contestable.

Journal: Proceedings of the 43rd International Conference on Machine Learning (ICML 2026), 2026

arXiv abs page · PDF · same-day batch