GAN FOR DATA SYNTHESIS WHILE MINIMIZING PREDICTIVE POWER LOSS
DOI:
https://doi.org/10.28925/2663-4023.2026.34.1241Keywords:
anonymization, financial risks, machine learning, confidentiality, data quality, fraudAbstract
In the modern financial environment, the ability to leverage data for developing intelligent risk detection methods is a huge ethical challenge. However, this does not make it less of a necessity. That is something I need to solve for my research as well. The quality and quantity of the data I would be able to employ for my research hinges on how good would be my anonymization proposal. Anonymization is the process of transforming datasets to prevent identification of individuals or organizations. It has become of paramount importance in order to ensure privacy and compliance in experiments involving sensitive financial information. As regulatory frameworks such as the General Data Protection Regulation (GDPR) in the European Union and similar local laws worldwide are on the rise, ethical approaches to anonymization are no longer optional but essential to achieve legal compliance and to conduct responsible research. Advances in machine learning and big data analytics have increased the likelihood of re-identification, especially when anonymized data is combined with external datasets [1]. This raises critical questions about whether traditional anonymization techniques remain sufficient in high-risk financial contexts and can be still used to full extent. However, it is known that over-anonymization will degrade data quality, sometimes severely, limiting the ability of machine learning algorithms to detect necessary patterns in order to locate fraudulent transactions efficiently. Ethical practice requires balancing privacy protections with the need for accurate and efficient models. Finding this balance is proving to be even more challenging as the demand for the data gets bigger. Removing key values will unintentionally bias models, particularly in fraud detection, potentially worsening their performance as well as their interpretability.
Downloads
References
Ohm, P. (2010). Broken promises of privacy: Responding to the surprising failure of anonymization. UCLA Law Review, 57(6), 1701-1777.
Kim, K. (2025). Self-refining language model anonymizers via adversarial distillation.
Rocha, Á. (n.d.). Generative AI in FinTech: Revolutionizing finance through intelligent algorithms.
Kaplan, J. (2025). Generative artificial intelligence: What everyone needs to know. Oxford University Press.
Goodfellow, I. J., et al. (2014). Generative adversarial nets. Advances in Neural Information Processing Systems, 27.
Xu, L., et al. (2019). Modeling tabular data using conditional GAN. Advances in Neural Information Processing Systems.
Arjovsky, M., Chintala, S., & Bottou, L. (2017). Wasserstein generative adversarial networks. In Proceedings of the 34th International Conference on Machine Learning (pp. 214-223).
Alqulaity, M., & Yang, P. (2024). Enhanced conditional GAN for high-quality synthetic tabular data generation in mobile-based cardiovascular healthcare. Sensors, 24(23), 7673.
Engelmann, J., & Lessmann, S. (2020). Conditional Wasserstein GAN-based oversampling of tabular data for imbalanced learning.
Radford, A., Metz, L., & Chintala, S. (2016). Unsupervised representation learning with deep convolutional generative adversarial networks.
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Дмитро Масюк, Георгій Гайна

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.