Logistic Regression with Incompletely Observed Categorical Covariates: A Comparison of Three Approaches
比较了处理分类协变量缺失的三种方法:最大似然估计、伪最大似然估计和概率插补,发现前两者一致且渐近正态,而条件概率插补接近伪最大似然,无条件概率插补有严重偏误。
A logistic regression analysis based on the complete cases neglects the information due to subjects with missing values in at least one but not in all covariates. Three different approaches allowing the use of this information are compared: maximum likelihood estimation, pseudo maximum likelihood estimation, and probability imputation. The first two yield consistent estimates of the regression coefficients and asymptotic normality allows the construction of asymptotically valid confidence intervals. The only necessary assumption is that the probability for the occurrence of missing values does not depend on the true value of the hidden covariates, whereas there may be a dependence on completely observed covariates and on the outcome variable. Their relative efficiency and the gain relative to a complete case analysis is investigated. Comparison with imputation methods shows that imputation of conditional probabilities can be regarded as a close approximation to the pseudo maximum likelihood estimation, whereas imputation of unconditional probabilities is associated with serious bias.