去随机化Knockoffs:利用e值控制错误发现率

Derandomised knockoffs: leveraging e-values for false discovery rate control

Journal of the Royal Statistical Society. Series B: Statistical Methodology · 2023
被引 31 · 同刊同年前 4%
ABS 4

中文导读

提出一种去随机化模型X knockoffs的方法,通过聚合多次knockoffs的e值来保证错误发现率控制,同时降低选择变量的变异性。

Abstract

Abstract Model-X knockoffs is a flexible wrapper method for high-dimensional regression algorithms, which provides guaranteed control of the false discovery rate (FDR). Due to the randomness inherent to the method, different runs of model-X knockoffs on the same dataset often result in different sets of selected variables, which is undesirable in practice. In this article, we introduce a methodology for derandomising model-X knockoffs with provable FDR control. The key insight of our proposed method lies in the discovery that the knockoffs procedure is in essence an e-BH procedure. We make use of this connection and derandomise model-X knockoffs by aggregating the e-values resulting from multiple knockoff realisations. We prove that the derandomised procedure controls the FDR at the desired level, without any additional conditions (in contrast, previously proposed methods for derandomisation are not able to guarantee FDR control). The proposed method is evaluated with numerical experiments, where we find that the derandomised procedure achieves comparable power and dramatically decreased selection variability when compared with model-X knockoffs.

高维回归错误发现率控制统计推断机器学习