联邦特征选择与错误发现率控制

Federated feature selection with false discovery rate control

Journal of the Royal Statistical Society. Series B: Statistical Methodology · 2025
被引 0
ABS 4

中文导读

提出Fed-FDR框架,在多个数据站点间选择与响应变量普遍相关的特征,同时控制错误发现率,仅共享低维系数估计以保护隐私,对特征分布和模型参数的异质性具有鲁棒性。

Abstract

Abstract Selecting a set of universally relevant features associated with a given response variable across multiple distributed data sites is an important problem in numerous scientific fields. However, performing this federated feature selection task becomes challenging when individual-level data cannot be shared due to privacy concerns. The problem is further complicated by potential heterogeneity in both feature distributions and model parameters across sites. In this paper, we propose Fed-false discovery rate (FDR), a federated feature selection framework that simultaneously identifies important features while controlling the FDR. To ensure privacy preservation and reduce communication costs, the Fed-FDR shares only lower-dimensional coefficient estimates instead of transmitting summary statistics for all features, with the dimensionality shown to be of the same order as the number of relevant features. The coordinating centre then leverages these lower-dimensional coefficient estimates to construct a generalized mirror statistic to identify the important features. The Fed-FDR is robust to the heterogeneity of feature distribution and model parameters, easy to implement, and computationally efficient. We further demonstrate that Fed-FDR effectively controls the FDR while achieving strong statistical power in our simulation studies. The results of the empirical study also demonstrate that the method is both valid and implementation-ready.

联邦学习特征选择错误发现率控制隐私保护