缓解信用评分中Transformer模型面临的对抗攻击

Mitigating adversarial attacks on transformer models in credit scoring

European Journal of Operational Research · 2025
被引 4
ABS 4

中文导读

研究了使用借款人文本数据的Transformer信用评分模型易受对抗攻击的脆弱性,并评估了对抗训练和主题建模两种缓解策略,发现它们能提升模型稳健性、减少定价错误,对金融从业者和监管者有参考价值。

Abstract

The integration of unstructured data, such as text created by borrowers, offers new opportunities for improving credit default prediction but also introduces new risks. This study examines the robustness of transformer-based credit scoring models that utilize textual data and assesses their vulnerability to adversarial attacks. Using peer-to-peer lending data, we show that small, semantically neutral changes in loan descriptions can substantially alter model outputs. These vulnerabilities expose lenders and borrowers to economic risks through distorted risk assessments and mispriced loans. We evaluate two mitigation strategies: adversarial training and topic modeling. Adversarial training improves robustness without compromising predictive performance. Topic modeling provides a more interpretable and stable representation of borrower narratives. An economic analysis confirms that robust models reduce mispricing and improve outcomes for all parties. The findings underscore the importance of robustness as the use of unstructured data in credit scoring becomes more accessible.

信用评分对抗攻击Transformer模型金融科技风险管理