Detection of financial fraud and recidivism using machine learning algorithms: UK evidence
本文利用英国金融监管机构2010至2022年起诉案例数据,比较多种机器学习算法,发现随机森林在预测财务欺诈和再犯方面表现最优,并识别出杠杆、公司规模等关键因素,为会计师和监管者提供早期预警参考。
This paper explores the effectiveness of various machine learning algorithms in predicting financial fraud and recidivism using a hand-collected dataset of cases prosecuted by UK financial regulators from 2010 to 2022. The study aims to identify key factors for predicting financial fraud and recidivism. Comparative analysis of machine learning algorithms, including Logistic Regression, Ridge Regression, Support Vector Machine, Decision Tree, Random Forest, Artificial Neural Network, and Adaptive Least Absolute Shrinkage and Selection Operator, reveals that the Random Forest model consistently outperforms others in AUC and precision rate for both financial fraud and recidivism. The study identifies crucial factors, including financial factors such as leverage, firm size, market value, and tangibility, and non-financial factors such as firm age, tone, gender ratio, and the number of directors, contributing significantly to financial fraud detection. The insights provide valuable guidance to accountants, independent directors and regulators for developing effective early warning systems for financial fraud.