A feature selection algorithm optimizing fitting and predictive performance of logistic regression: a case study on financial literacy and pension planning
提出一种联合优化逻辑回归解释与预测能力的特征选择算法,通过前向搜索迭代选取能显著提升AUC的变量,并以意大利银行2020年家庭收入与财富调查数据为例,分析金融素养对养老金规划的影响。
Abstract When dealing with binary regression problems, it is common practice to address first variable selection and model fitting on a training set, and then to assess its prediction performance on a test set. In this setting, the paper casts a feature selection algorithm for logistic regression that jointly optimizes explicative and predictive abilities of the available information set. To this aim, a forward search is implemented within the covariate space that iteratively selects the predictor whose inclusion in the model yields the highest significant increase in the Area Under the ROC curve (AUC) with respect to the previous step. The resulting procedure adheres to a parsimony principle and returns the relative contribution of each regressor in the prediction accuracy of the final model. The proposal is show-cased with a study on financial literacy and pension planning, on the wake of the survey on Household Income and Wealth run by the Bank of Italy in 2020. Indeed, recent literature in behavioral economic and finance highlight that boosting financial literacy among the population is a key strategy to sustain efficacy of both public policies and individual well-being face to the ageing of the population and times of economic crises.