PaperScope
LIVE · 2026-09-29 05:40 UTC

Collaborative Synthetic Data for Privacy-Preserving Financial Fraud Detection Across Organizational Silos

Simeon Allmendinger, Domenique Zipperling, Burhanettin Bahadir Kibar, Niklas K{ü}hl

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.33754 v1
Category
Submitted
2026-09-27

Abstract

Organizations seek analytical value from AI, yet relevant data are often fragmented across organizations and constrained by privacy. This is acute in financial fraud detection, where rare fraud cases and imbalanced local datasets limit decision-relevant analytics. Federated learning enables collaboration without direct data sharing but does not resolve minority-class scarcity. Synthetic data generation can help, yet lightweight methods are interpolation-bound, while generative models require substantial data and computation. Existing collaborative generative approaches often rely on federated learning, imposing considerable organization-side training burdens. In this paper, we examine CollaFuse as a collaborative diffusion-based alternative for fraud detection and evaluate it across five fraud datasets. Compared with classical oversampling, local generative baselines, and centralized diffusion benchmarks, CollaFuse does not achieve the highest local fidelity but improves downstream fraud detection more consistently across most datasets. These findings suggest that synthetic data create analytical value less through local realism than through transferable cross-organizational structure.

Comment: Accepted for publication at the International Conference on Information Systems (ICIS) 2026

arXiv abs page · PDF · same-day batch