Spectrum-aware debiasing: A modern inference framework with applications to principal components regression
提出一种谱感知去偏方法,适用于具有结构化行-列依赖、重尾、非对称和潜在低秩结构的高维回归问题,实现了主成分回归中的去偏估计,并提供了渐近正态性和谱普适性保证。
Debiasing is a fundamental concept in high-dimensional statistics. While degrees-of-freedom adjustment is the state-of-the-art debiasing technique in high-dimensional linear regression, it largely remains limited to independent, identically distributed samples and sub-Gaussian covariates. These limitations hinder its wider practical use. In this paper we break this barrier and introduce Spectrum-Aware Debiasing—a novel inference method that applies to challenging high-dimensional regression problems with structured row-column dependencies, heavy tails, asymmetric properties, and latent low-rank structures. Our method achieves debiasing through a rescaled gradient descent step, where the rescaling factor is derived from the spectral properties of the sample covariance matrix. This spectrum-based approach enables accurate debiasing in much broader contexts. We study the common modern regime where the number of features and samples scale proportionally. We establish asymptotic normality of our proposed estimator (suitably centered and scaled) under various convergence notions when the covariates are right-rotationally invariant. We further prove a spectral universality result, extending our guarantees to a much broader class of covariate distributions. Furthermore, we devise a consistent estimator for the asymptotic variance. Our work has two notable by-products: First, Spectrum-Aware Debiasing rectifies the bias in principal components regression (PCR), providing the first debiased PCR estimator in high dimensions. Second, we introduce a principled test for checking the presence of alignment between the signal and the eigenvectors of the sample covariance matrix. This test is independently valuable for statistical methods developed using approximate message passing, leave-one-out, random matrix theory, or convex Gaussian min-max theorems. We demonstrate the utility of our method through diverse simulated and real data experiments.