Fighting sampling bias: A framework for training and evaluating credit scoring models
研究了信用评分模型中因忽略被拒申请人导致的抽样偏差问题,提出了偏差感知自标记算法和贝叶斯评估框架,实验表明新方法在预测准确性和盈利能力上优于现有基准。
Scoring models support decision-making in financial institutions. Their estimation and evaluation rely on labeled data from previously accepted clients. Ignoring rejected applicants with unknown repayment behavior introduces sampling bias, as the available labeled data only partially represents the population of potential borrowers. This paper examines the impact of sampling bias and introduces new methods to mitigate its adverse effect. First, we develop a bias-aware self-labeling algorithm for scorecard training, which debiases the training data by adding selected rejects with an inferred label. Second, we propose a Bayesian framework to address sampling bias in scorecard evaluation. To provide reliable projections of future scorecard performance, we include rejected clients with random pseudo-labels in the test set and use Monte Carlo sampling to estimate the scorecard’s expected performance across label realizations. We conduct extensive experiments using both synthetic and observational data. The observational data includes an unbiased sample of applicants accepted without scoring, representing the true borrower population and facilitating a realistic assessment of reject inference techniques. The results show that our methods outperform established benchmarks in predictive accuracy and profitability. Additional sensitivity analysis clarifies the conditions under which they are most effective. Comparing the relative effectiveness of addressing sampling bias during scorecard training versus evaluation, we find the latter much more promising. For example, we estimate the expected return per dollar issued to increase by up to 2.07 and up to 5.76 percentage points when using bias-aware self-labeling and Bayesian evaluation, respectively. • New reject inference methods for scorecard training and evaluation. • Real micro-lending data without bias uncovers the true impact of reject inference. • Sensitivity analysis and simulations clarify the boundary conditions of the new methods. • Bias-aware evaluation has a higher business impact than labeling rejects for scorecard training.