基于插件抽样生成的合成数据的多元回归模型推断

Inference for Multivariate Regression Model Based on Synthetic Data Generated Using Plug-in Sampling

Journal of the American Statistical Association · 2021
被引 7
ABS 4

中文导读

本文推导了多元回归模型中基于插件抽样生成的合成数据的似然精确推断方法,并通过模拟和实际数据验证其有效性,对处理隐私保护数据的统计学者有用。

Abstract

In this article, the authors derive the likelihood-based exact inference for singly and multiply imputed synthetic data in the context of a multivariate regression model. The synthetic data are generated via the Plug-in Sampling method, where the unknown parameters in the model are set equal to the observed values of their point estimators based on the original data, and synthetic data are drawn from this estimated version of the model. Simulation studies are carried out in order to confirm the theoretical results. The authors provide exact test procedures, which in case multiple synthetic datasets are permissible, are compared with the asymptotic results of Reiter. An application using 2000 U.S. Current Population Survey public use data is discussed. Furthermore, properties of the proposed methodology are evaluated in scenarios where some of the conditions that were used to derive the methodology do not hold, namely for nonnormal and discrete distributed random variables, cases in which the inferential procedures developed still show very good performances.

多元统计合成数据回归分析统计推断数据隐私