PaperScope
LIVE · 2026-09-29 05:40 UTC

Fuzzy Distribution Modeling for Synthetic Tabular Data Generation with Causality Preservation

Michael Vasilakakis, Dimitris K. Iakovidis

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.34349 v1
Category
Submitted
2026-09-28

Abstract

Synthetic tabular data generation provides an effective alternative for the training of machine learning models when real-world data is limited or inaccessible. However, the heterogeneous, non-smooth, and incomplete nature of tabular data poses fundamental challenges to conventional probabilistic and deep generative models, where their interpretability remains limited. This paper proposes a novel fuzzy distribution modeling methodology for synthetic tabular data generation based on fuzzy sets theory. Feature distributions are represented using fuzzy sets and feature dependencies are modeled through Fuzzy Cognitive Maps, resulting in a low-parameter, and an interpretable data representation. Synthetic samples are generated by sampling fuzzy concepts rather than raw values, enabling native support for mixed data types, missing values, and domain constraints. The methodology further supports linguistic queries and IF-THEN reasoning, facilitating transparent simulation of decision-making processes. Experimental results on benchmark datasets demonstrate competitive performance with respect to utility, fidelity and privacy compared to state-of-the-art methods, while offering substantially improved interpretability. These results establish fuzzy distribution modeling as a principled and effective approach for synthetic tabular data generation in fuzzy systems and decision support applications.

Comment: 6 pages, 2 figures, 3 tables. Published in the 2026 IEEE International Conference on Fuzzy Systems (FUZZ-IEEE), Maastricht, the Netherlands

Journal: Proc. 2026 IEEE Int. Conf. on Fuzzy Systems (FUZZ-IEEE), pp. 1-6

arXiv abs page · PDF · same-day batch