用于信用卡欺诈检测中数据增强和迁移的生成对抗网络

Generative adversarial networks for data augmentation and transfer in credit card fraud detection

Journal of the Operational Research Society · 2021
被引 45 · 同刊同年前 3%
ABS 3

中文导读

研究了在信用卡欺诈检测中,使用生成对抗网络生成合成样本进行数据增强,并考虑客户信用质量分布,发现高信用质量客户机构更受益于增强,而低信用质量银行在数据迁移中获益更多。

Abstract

Augmenting a dataset with synthetic samples is a common processing step in machine learning with imbalanced classes to improve model performance. Another potential benefit of synthetic data is the ability to share information between cooperating parties while maintaining customer privacy. Often overlooked, however, is how the distribution of the data affects the potential gains from synthetic data augmentation. We present a case study in credit card fraud detection using Generative Adversarial Networks to generate synthetic samples, with explicit consideration given to customer distributions. We investigate two different cooperating party scenarios yielding four distinct customer distributions by credit quality. Our findings indicate that institutions skewed towards higher credit quality customers are more likely to benefit from augmentation with GANs. Relative gains from synthetic data transfer, in the absence of feature set heterogeneity, also appear to asymmetrically favour banks operating on the lower end of the credit spectrum, which we hypothesise is due to differences in spending behaviours.

机器学习信用卡欺诈检测生成对抗网络数据增强迁移学习