AI Techniques for Survey Data Quality: Transformers and GANs
By: Simona Cafieri and Gianmarco Borrata
References
- Aittokallio, T. (2010). Dealing with missing values in large-scale studies: Microarray data imputation and beyond. Briefings in Bioinformatics, 11(2), 253–264. https://doi.org/10.1093/bib/bbp060
- Baraldi, A. N., & Enders, C. K. (2010). An introduction to modern missing data analyses. Journal of School Psychology, 48(1), 5–37. https://doi.org/10.1016/j.jsp.2009.10.001
- Belgiu, M., & Drăguț, L. (2016). Random forest in remote sensing: A review of applications and future directions. ISPRS Journal of Photogrammetry and Remote Sensing, 114, 24–31. https://doi.org/10.1016/j.isprsjprs.2016.01.011
- Casella, M., Milano, N., Dolce, P., & Marocco, D. (2024). Transformers deep learning models for missing data imputation: An application of the ReMasker model on a psychometric scale. Frontiers in Psychology, 15, Article 1449272. https://doi.org/10.3389/fpsyg.2024.1449272
- De Leeuw, E. D. (2011). Reducing missing data in surveys: An overview of methods. Quality & Quantity, 45(1), 147–160. https://doi.org/10.1007/s11135-010-9350-9
- Dempster, A. P., Laird, N. M., & Rubin, D. B. (1977). Maximum likelihood from incomplete data via the EM algorithm. Journal of the Royal Statistical Society: Series B (Methodological), 39(1), 1–38. https://doi.org/10.1111/j.2517-6161.1977.tb01600.x
- Dey, R., & Salem, F. M. (2017). Gate-variants of gated recurrent unit (GRU) neural networks. In 2017 IEEE 60th International Midwest Symposium on Circuits and Systems (MWSCAS) (pp. 1597–1600). IEEE. https://doi.org/10.1109/MWSCAS.2017.8053243
- Du, T., Melis, L., & Wang, T. (2023). ReMasker: Imputing tabular data with masked autoencoding. arXiv. https://arxiv.org/abs/2309.13793
- García-Laencina, P. J., Sancho-Gómez, J. L., & Figueiras-Vidal, A. R. (2010). Pattern classification with missing data: A review. Neural Computing and Applications, 19, 263–282. https://doi.org/10.1007/s00521-009-0295-6
- Gondara, L., & Wang, K. (2017). Multiple imputation using deep denoising autoencoders. arXiv. https://arxiv.org/abs/1705.02737
- Graham, J. W., Olchowski, A. E., & Gilreath, T. D. (2007). A gentle introduction to imputation of missing values. Prevention Science, 8(3), 206–213. https://doi.org/10.1007/s11121-007-0099-9
- Gumbel, E. J. (1954). Statistical theory of extreme values and some practical applications. National Bureau of Standards.
- Harel, O., & Zhou, X.-H. (2007). Multiple imputation: Review of theory, implementation, and software. Statistics in Medicine, 26(16), 3057–3077. https://doi.org/10.1002/sim.2787
- Hochreiter, S., & Schmidhuber, J. (1997). Long short-term memory. Neural Computation, 9(8), 1735–1780. https://doi.org/10.1162/neco.1997.9.8.1735
- ISTAT. (2022). Indagine Aspetti della vita quotidiana 2021 [Data set]. Istituto Nazionale di Statistica.
- ISTAT. (2023). Indagine Aspetti della vita quotidiana 2022 [Data set]. Istituto Nazionale di Statistica.
- Jang, E., Gu, S., & Poole, B. (2017). Categorical reparameterization with Gumbel-Softmax. In International Conference on Learning Representations.
- Liew, A. W.-C., Law, N.-F., & Yan, H. (2011). Missing value imputation for gene expression data: Computational techniques to recover missing data from available information. Briefings in Bioinformatics, 12(5), 498–513. https://doi.org/10.1093/bib/bbq080
- Lin, W.-C., & Tsai, C.-F. (2020). Missing value imputation: A review and analysis of the literature. Artificial Intelligence Review, 53, 1487–1509. https://doi.org/10.1007/s10462-019-09709-4
- Little, R. J. A., & Rubin, D. B. (2019). Statistical analysis with missing data (3rd ed.). Wiley.
- Nikfalazar, S., Yeh, C.-H., Bedingfield, S., & Khorshidi, H. A. (2020). Missing data imputation using decision trees and fuzzy clustering with iterative learning. Knowledge and Information Systems, 62, 2419–2437. https://doi.org/10.1007/s10115-020-01469-x
- Rubin, D. B. (1976). Inference and missing data. Biometrika, 63(3), 581–592. https://doi.org/10.1093/biomet/63.3.581
- Rubin, D. B. (1987). Multiple imputation for nonresponse in surveys. Wiley.
- Sun, Y., Li, J., Xu, Y., Zhang, T., & Wang, X. (2023). Deep learning versus conventional methods for missing data imputation: A review and comparative study. Expert Systems with Applications, 227, Article 120201. https://doi.org/10.1016/j.eswa.2023.120201
- Tsikriktsis, N. (2005). A review of techniques for treating missing data in OM survey research. Journal of Operations Management, 24(1), 53–62. https://doi.org/10.1016/j.jom.2005.03.001
- Yoon, J., Jordon, J., & van der Schaar, M. (2018). GAIN: Missing data imputation using generative adversarial nets. In Proceedings of the 35th International Conference on Machine Learning (pp. 5689–5698). PMLR.
DOI: https://doi.org/10.2478/jses-2026-0005 | Journal eISSN: 2285-388X
Language: English
Page range: 57 - 68
Published on: Jun 29, 2026
Published by: Bucharest University of Economic Studies
In partnership with: Paradigm Publishing Services
Publication frequency: 2 issues per year
Related subjects:
© 2026 Simona Cafieri, Gianmarco Borrata, published by Bucharest University of Economic Studies
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License.