Privacy through Synthesis: Exploration of Tabular Synthetic Data Generation Methods and Quality Assessment
By: Bhagyashree Chougule and Pooja Bagane
References
- I. Y. Jung, ‘A review of privacy-preserving human and human activity recognition’, Int. J. Smart Sens. Intell. Syst., vol. 13, no. 1, pp. 1–13, 2020.
- D. V. Kute, B. Pradhan, N. Shukla, and A. Alamri, ‘Explainable deep learning model for predicting money laundering transactions’, Int. J. Smart Sens. Intell. Syst., vol. 17, no. 1, 2024.
- H. Gandhi, K. Tandon, S. Gite, B. Pradhan, and A. Alamri, ‘Navigating the Complexity of Money Laundering: Anti-money Laundering Advancements with AI/ML Insights’, Int. J. Smart Sens. Intell. Syst., vol. 17, no. 1, 2024.
- Rubin D B, ‘Statistical disclosure limitation’, J. Off. Stat., vol. 9, no. 2, pp. 461–468, 1993.
- G. Fioretti, ‘Agent-based simulation models in organization science’, Organ. Res. methods, vol. 16, no. 2, pp. 227–242, 2013.
- I. J. Goodfellow et al., ‘Generative adversarial nets’, in Advances in Neural Information Processing Systems, 2014, pp. 2672–2680.
- D. P. Kingma and M. Welling, ‘Auto-encoding variational bayes’, in 2nd International Conference on Learning Representations, ICLR 2014 - Conference Track Proceedings, 2014.
- S. James, C. Harbron, J. Branson, and M. Sundler, ‘Synthetic data use: exploring use cases to optimise data utility’, Dec. 01, 2021, Springer Nature.
- D. Shamaev, ‘Synthetic Datasets and Medical Artificial Intelligence Specifics’, in Lecture Notes in Networks and Systems, Springer Science and Business Media Deutschland GmbH, 2023, pp. 519–528.
- M. Rigaki and S. Garcia, ‘A Survey of Privacy Attacks in Machine Learning’, ACM Comput. Surv., vol. 56, no. 4, Apr. 2024,
- R. Shokri, M. Stronati, … C. S.-2017 I. symposium, and undefined 2017, ‘Membership inference attacks against machine learning models’, ieeexplore. ieee.orgR Shokri, M Stronati, C Song, V Shmatikov2017 IEEE Symp. Secur. Priv. (SP), 2017•ieeexplore.ieee.org.
- S. Hisamoto, M. Post, and K. Duh, ‘Membership inference attacks on sequence-to-sequence models: Is my data in your machine translation system?’, direct.mit.eduS Hisamoto, M Post, K DuhTransactions Assoc. Comput. Linguist. 2020•direct.mit.edu,vol. 8, pp. 49–63, 2020,
- K. Shu, S. Wang, J. Tang, R. Zafarani, and H. Liu, ‘User Identity Linkage across Online Social Networks’, ACM SIGKDD Explor. Newsl., vol. 18, no. 2, pp. 5–17, Mar. 2017,
- N. Z. Gong and B. Liu, ‘You are who you know and how you behave: Attribute inference attacks via users’ social friends and behaviors’, in Proceedings of the 25th USENIX Security Symposium, 2016, pp. 979–995.
- A. Narayanan and V. Shmatikov, ‘Robust De-anonymization of Large Sparse Datasets’, in 2008 IEEE Symposium on Security and Privacy (sp 2008), 2008.
- S. Mahloujifar, E. Ghosh, and M. Chase, ‘Property Inference from Poisoning’, in Proceedings - IEEE Symposium on Security and Privacy, 2022, pp. 1120–1137.
- K. Ganju, Q. Wang, W. Yang, C. A. Gunter, and N. Borisov, ‘Property inference attacks on fully connected neural networks using permutation invariant representations’, in Proceedings of the ACM Conference on Computer and Communications Security, Association for Computing Machinery, Oct. 2018, pp. 619–633.
- G. Ateniese, L. V. Mancini, A. Spognardi, A. Villani, D. Vitali, and G. Felici, ‘Hacking smart machines with smarter ones: How to extract meaningful data from machine learning classifiers’, Int. J. Secur. Networks, vol. 10, no. 3, pp. 137–150, Sep. 2015,
- F. Guan, T. Zhu, H. Tong, W. Z.-K.-B. Systems, and U. 2024, ‘A realistic model extraction attack against graph neural networks’, ElsevierF Guan, T Zhu, H Tong, W ZhouKnowledge-Based Syst. 2024•Elsevier, 2024.
- M. Nasr, R. Shokri, and A. Houmansadr, ‘Comprehensive privacy analysis of deep learning: Passive and active white-box inference attacks against centralized and federated learning’, in Proceedings - IEEE Symposium on Security and Privacy, 2019, pp. 739–753.
- N. Papernot, P. McDaniel, I. Goodfellow, S. Jha, Z. B. Celik, and A. Swami, ‘Practical black-box attacks against machine learning’, dl.acm.orgN Pap. P McDaniel, I Goodfellow, S Jha, ZB Celik, A SwamiProceedings 2017 ACM Asia Conf. Comput. and, 2017•dl.acm.org, pp. 506–519, Apr. 2017,
- B. Wang and N. Z. Gong, ‘Stealing Hyperparameters in Machine Learning’, in Proceedings - IEEE Symposium on Security and Privacy, 2018, pp. 36–52.
- S. J. Oh, B. Schiele, and M. Fritz, ‘Towards Reverse-Engineering Black-Box Neural Networks’, in Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), vol. 11700 LNCS, Springer Verlag, 2019, pp. 121–144.
- B. Jayaraman and D. Evans, ‘Evaluating differentially private machine learning in practice’, in Proceedings of the 28th USENIX Security Symposium, 2019, pp. 1895–1912.
- L. H. Newman, ‘Newman, L. H. (2018). A new Google+ blunder exposed… - Google Scholar’.
- M. Yang, T. Guo, T. Zhu, I. Tjuawinata, … J. Z.-C. S. &, and U. 2024, ‘Local differential privacy and its applications: A comprehensive survey’, Elsevier, 2024.
- K. El Emam, L. Mosquera, and R. Hoptroff, Practical synthetic data generation: balancing privacy and the broad availability of data. O’Reilly Media, 2020.
- S. Nikolenko, Synthetic data for deep learning. Springer, 2021.
- M. Goyal, Q. M.- Electronics, and U. 2024, ‘A systematic review of synthetic data generation techniques using generative AI’, mdpi.comM Goyal, QH MahmoudElectronics, 2024•mdpi.com, 2024.
- F. M. Hollenbach, I. Bojinov, S. Minhas, N. W. Metternich, M. D. Ward, and A. Volfovsky, ‘Multiple Imputation Using Gaussian Copulas’, Sociol. Methods Res., vol. 50, no. 3, pp. 1259–1283, Aug. 2021,
- R. Batuwita and V. Palade, ‘Efficient resampling methods for training support vector machines with imbalanced datasets’, in Proceedings of the International Joint Conference on Neural Networks, 2010.
- N. Chawla, K. Bowyer, … L. H.-J. of artificial, and U. 2002, ‘SMOTE: synthetic minority over-sampling technique’, jair.orgNV Chawla, KW Bowyer, LO Hall, WP KegelmeyerJournal Artif. Intell. Res. 2002•jair.org, vol. 16, pp. 321–357, 2002.
- A. Figueira and B. Vaz, ‘Survey on Synthetic Data Generation, Evaluation Methods and GANs’, Mathematics, vol. 10, no. 15, pp. 1–41, 2022,
- H. Han, W. Y. Wang, and B. H. Mao, ‘Borderline-SMOTE: A new over-sampling method in imbalanced data sets learning’, in Lecture Notes in Computer Science, Springer Verlag, 2005, pp. 878–887.
- C. Bunkhumpornpat, K. Sinapiromsaran, and C. Lursinsap, ‘Safe-level-SMOTE: Safe-level-synthetic minority over-sampling technique for handling the class imbalanced problem’, in Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 2009, pp. 475–482.
- H. He, Y. Bai, E. Garcia, S. L.-2008 I. international Joint, and U. 2008, ‘ADASYN: Adaptive synthetic sampling approach for imbalanced learning’, ieeexplore.ieee.orgH He, Y Bai, EA Garcia, S Li2008 IEEE Int. Jt. Conf. neural networks (IEEE, 2008•ieeexplore. ieee.org, 2008.
- G. Douzas, F. Bacao, and F. Last, ‘Improving imbalanced learning through a heuristic oversampling method based on k-means and SMOTE’, Inf. Sci. (Ny)., vol. 465, pp. 1–20, 2018,
- T. Jo and N. Japkowicz, ‘Class imbalances versus small disjuncts’, ACM SIGKDD Explor. Newsl., vol. 6, no. 1, pp. 40–49, Jun. 2004,
- F. Figueira, Á., & Renna, New Insights in Machine Learning and Deep Neural Networks. MDPI-Multidisciplinary Digital Publishing Institute., 2023.
- D. Scikit-learn, ‘Gaussian Mixture Models’, scikit-learn.org. [Online]. Available:
https://scikit-learn.org/stable/modules/mixture.html - David Foster, Generative Deep Learning 2nd Edition. 2023.
- Z. Islam, M. Abdel-Aty, Q. Cai, and J. Yuan, ‘Crash data augmentation using variational autoencoder’, Accid. Anal. Prev., vol. 151, 2021,
- L. Xu and K. Veeramachaneni, ‘Synthesizing Tabular Data using Generative Adversarial Networks’, Nov. 2018.
- L. Xu, M. Skoularidou, A. Cuesta-Infante, and K. Veeramachaneni, ‘Modeling tabular data using conditional GAN’, in Advances in Neural Information Processing Systems, 2019.
- A. Rajabi and O. O. Garibay, ‘TabFairGAN: Fair Tabular Data Generation with Generative Adversarial Networks’, Mach. Learn. Knowl. Extr., vol. 4, no. 2, pp. 488–501, 2022
- F. J. Anscombe, ‘Graphs in statistical analysis’, Am. Stat., vol. 27, no. 1, pp. 17–21, 1973,
- Y. Zhang et al., ‘On the Properties of Kullback-Leibler Divergence Between Multivariate Gaussian Distributions’, in Advances in Neural Information Processing Systems, 2023.
- A. Stéphanovitch, U. Tanielian, B. Cadre, N. Klutchnikoff, and G. Biau, ‘Optimal 1-Wasserstein distance for WGANs’, Bernoulli, vol. 30, no. 4, pp. 2955–2978, Nov. 2024,
- F. Ji, X. Zhang, and J. Zhao, ‘α-EGAN: α-Energy distance GAN with an early stopping rule’, Comput. Vis. Image Underst., vol. 234, 2023,
- H. Gao and X. Shao, ‘Two Sample Testing in High Dimension via Maximum Mean Discrepancy’, J. Mach. Learn. Res., vol. 24, pp. 1–33, 2023.
- A. Goncalves, P. Ray, B. Soper, J. Stevens, L. Coyle, and A. P. Sales, ‘Generation and evaluation of synthetic patient data’, BMC Med. Res. Methodol., vol. 20, no. 1, pp. 1–40, 2020
- S. Liaskovska, S. Tyskyi, Y. Martyn, A. T. Augousti, and V. Kulyk, ‘Systematic Generation and Evaluation of Synthetic Production Data for Industry 5.0 Optimization’, Technologies, vol. 13, no. 2, 2025
- M. Miletic and M. Sariyar, ‘Challenges of Using Synthetic Data Generation Methods for Tabular Microdata’, Appl. Sci., vol. 14, no. 14, p. 5975, Jul. 2024
- K. Fang, V. Mugunthan, V. Ramkumar, and L. Kagal, ‘Overcoming Challenges of Synthetic Data Generation’, in Proceedings - 2022 IEEE International Conference on Big Data, Big Data 2022, Institute of Electrical and Electronics Engineers Inc., 2022, pp. 262–270.
- J. Drechsler, ‘Challenges in Measuring Utility for Fully Synthetic Data’, in Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), Springer Science and Business Media Deutschland GmbH, 2022, pp. 220–233.
- Z. Zhang, C. Yan, and B. A. Malin, ‘Membership inference attacks against synthetic health data’, J. Biomed. Inform., vol. 125, p. 103977, Jan. 2022,
DOI: https://doi.org/10.2478/ijssis-2026-0040 | Journal eISSN: 1178-5608
Language: English
Submitted on: Jun 20, 2025
Published on: Jul 11, 2026
Published by: International Journal on Smart Sensing and Intelligent Systems
In partnership with: Paradigm Publishing Services
Publication frequency: 1 issue per year
Related subjects:
© 2026 Bhagyashree Chougule, Pooja Bagane, published by International Journal on Smart Sensing and Intelligent Systems
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License.