Outliers and Residual Distributions in Logistic Regression
本文指出线性回归的残差诊断方法不能直接套用到逻辑回归,因为逻辑回归的残差分布依赖于解释变量,且小概率事件对估计有特殊影响。
Abstract Detection of outliers and other diagnostics based on residuals have gained widespread use in linear regression. Logistic regression has been both blessed and hindered by this development. Certainly logistic regression requires procedures to detect global and local model weaknesses. Thus the wealth of work done in linear regression provides guides and suggestions that may, with care and ingenuity, be applied to logistic regression. Several innovative attempts in this direction have been made by authors such as Tsiatis (1980), Pregibon (1979, 1981), Landwehr, Pregibon, and Shoemaker (1984), and Cook and Weisberg (1982). Unfortunately the similarities that allow such techniques to adapt to logistic regression seem, in addition, to hide many of the differences. Two such differences are discussed here. The first examines the effect of events of small probability on estimation in logistic regression. The simple conclusion is that to estimate the probability of success p in an area in which p is small, one must observe events of small probability. This suggests that the motives for outlier detection should be carefully considered. The second difference examined deals with the distribution of residuals. Let y be the response variable and let x be a p×1 vector of explanatory variable. In linear regression with normal errors, Y — E (Y | x) is normally distributed with mean 0 and variance σ2. Thus the distribution of this residual from the true model does not depend on the explanatory variables x. This is not true in logistic regression with binary data. There Y has a Bernoulli distribution with probability of success p(x) = Pr(Y = 1 | x). Thus any reasonable definition of a residual from the true model, for example, Y — p(x), has a two-point distribution that depends on x through p(x). That is, each residual has a unique distribution. In this setting claims that "standardized residuals" exist that are asymptotically normal or chi squared are not appropriate. The argument based on difference in likelihoods to support such claims as made by Pregibon (1981) and alluded to by Beckman and Cook (1983) relies on asymptotic conditions that are not obtained in strictly binary data. Diagnostics must consider the effects of differing residual distributions. For example, the local mean deviance plot suggested by Landwehr et al. (1984) is seriously affected by the underlying p(x i ). A modification, however, that replaces the deviance residual by (y — p)2 appears to be useful. The general intent of this article is to point out that direct application of linear regression techniques to logistic regression does not necessarily produce useful diagnostic tools. Thus careful consideration of each diagnostic technique is necessary. In addition, logistic regression might benefit from views and methods different from those applied to problems with continuous errors due to the 0–1 nature of the data. Key Words: Binary dataGoodness-of-fit testsResidual analysisDeviance