Estimating the number of significant components in high-dimensional principal component analysis
提出一种基于解释方差比和非尖峰样本特征值刚性的惩罚方法,用于估计高维主成分分析中的显著成分数量,在独立数据和部分时间序列数据下均具有一致性,且条件弱于AIC和BIC。
Abstract We consider the problem of estimating the number of significant components in high-dimensional principal component analysis. We propose a new penalized approach using the explained variance ratio and the rigidity of the nonspiked sample eigenvalues of sample covariance matrices of $ p $ variables. Compared with methods in the existing literature, the consistency of the proposed estimator holds, not only for independent data, but also for some times series data when the dimension $ p $ and the sample size $ n $ both tend to infinity. Even for independent data our estimator works under weaker conditions than existing approaches such as the aic and bic, including allowing heterogeneity in the bulk of the population eigenvalues. Simulation studies are conducted to illustrate the performance of the proposed estimator.