Small area estimation with generalized random forests: estimating poverty rates in Mexico
提出广义混合效应随机森林方法,用于在数据稀缺的小区域内估计贫困指标,处理非线性关系和层级结构。结合参数自助法评估误差,并在墨西哥Tlaxcala州展示贫困空间模式,适合研究贫困测度和小区域统计的学者参考。
Abstract Identifying and addressing poverty is challenging in administrative units with limited information on income distribution and well-being. To overcome this obstacle, small area estimation methods have been developed to provide reliable and efficient estimators at disaggregated levels, enabling informed decision-making by policymakers despite the data scarcity. We present the generalized mixed effects random forest, a robust and flexible framework for estimating area-level indicators under generalized modelling assumptions within the small area estimation context. The methodology accommodates response variables from the exponential family through a general link function; the empirical analyses in this article focus on poverty indicators based on binary response variables. Our method employs machine learning techniques to identify predictive, non-linear relationships from data, while also modelling hierarchical structures. Mean squared error estimation is explored using a parametric bootstrap. From an applied perspective, we examine the impact of information loss due to converting continuous variables into binary variables on the performance of small area estimation methods. We evaluate the proposed point and uncertainty estimates in both model- and design-based simulations. Finally, we apply our method to a case study revealing spatial patterns of poverty in the Mexican state of Tlaxcala.