分组数据部分线性模型中的三明治提升法用于精确估计

Sandwich boosting for accurate estimation in partially linear models for grouped data

Journal of the Royal Statistical Society. Series B: Statistical Methodology · 2024
被引 1
ABS 4

中文导读

针对分组数据部分线性模型,提出一种新损失函数(三明治损失)和梯度提升算法(三明治提升),在组内相关结构误设时仍能提升线性参数估计精度,适用于经济学等领域的聚类数据分析。

Abstract

Abstract We study partially linear models in settings where observations are arranged in independent groups but may exhibit within-group dependence. Existing approaches estimate linear model parameters through weighted least squares, with optimal weights (given by the inverse covariance of the response, conditional on the covariates) typically estimated by maximizing a (restricted) likelihood from random effects modelling or by using generalized estimating equations. We introduce a new ‘sandwich loss’ whose population minimizer coincides with the weights of these approaches when the parametric forms for the conditional covariance are well-specified, but can yield arbitrarily large improvements in linear parameter estimation accuracy when they are not. Under relatively mild conditions, our estimated coefficients are asymptotically Gaussian and enjoy minimal variance among estimators with weights restricted to a given class of functions, when user-chosen regression methods are used to estimate nuisance functions. We further expand the class of functional forms for the weights that may be fitted beyond parametric models by leveraging the flexibility of modern machine learning methods within a new gradient boosting scheme for minimizing the sandwich loss. We demonstrate the effectiveness of both the sandwich loss and what we call ‘sandwich boosting’ in a variety of settings with simulated and real-world data.

计量经济学机器学习统计估计分组数据