含分类协变量的非参数回归中的变量选择

Variable Selection in Nonparametric Regression with Categorical Covariates

Journal of the American Statistical Association · 1992
被引 11
ABS 4

中文导读

研究了在非参数回归模型中,当协变量为分类变量时如何选择预测变量。比较了交叉验证和累积预测误差两种准则,发现前者适用于无限维真实模型,后者适用于有限维模型,并通过模拟展示了小样本性质。

Abstract

Abstract This article extends the problem of variable selection to a nonparametric regression model with categorical covariates. Two selection criteria are considered: the cross-validation (CV) criterion and the accumulated prediction error (APE) criterion. We find that, asymptotically, the CV criterion performs well only when the true model is infinite-dimensional, while the APE criterion is appropriate when the true model is finite-dimensional. This is very similar to the case of linear regression model. A simulation study reveals some interesting small-sample properties of these criteria. To be more specific, suppose that we have observations (X 1, Y 1), …, (Xn, Yn ) that are iid random vectors and X = (X(1), X(2), …), where the X(i)'s are categorical. We allow Y to be of any type. Now a new observation X has arrived and we want to predict the corresponding Y. Such a framework is more appropriate than regressions with fixed covariates in situations where the covariates are observational rather than being controlled. For instance, Y could be the time from HIV infection to developing clinical AIDS, and the covariates (mostly categorical or reducible to categorical) could be observations from blood tests, a physical examination, or further personal information, such as sexual practices obtained from an interview. Take another example: Y could be the premium of an insurance policy with the covariates being the customer's general demographical information. Our goal is to select a subset of covariates that best predict Y. We define the true model dimension as d 0 if the regression function E(Y|X(1), X(2), …) is a d 0-variate function. The main conclusions of the article are: (1) The popular CV criterion performs well only when d 0 = ∞. (2) There exist other criteria that are more appropriate than CV when d 0 < ∞. (3) There is no difference between conditional and unconditional prediction errors, as far as asymptotics are concerned. (4) The selection range has to depend on the sample size. In fact, we argue that, for a given sample size n, we should only select models with the number of covariates not exceeding the order of magnitude of o(log n). (5) Simulation study indicates that the CV criterion has nice small-sample properties.

计量经济学非参数统计变量选择分类数据