Junhyung Chang and Xiaoyu Lei’s contribution to the Discussion of ‘Statistical exploration of the Manifold Hypothesis’ by Whiteley et al.
通过模拟研究,探讨随机函数选择如何影响谱聚类性能,并讨论流形假设下主成分分析的统计性质及假设检验方向。
We congratulate the authors on an insightful paper that spans both a general theory of manifold emergence, and methodology for exploring hidden manifold structure in high-dimensional data. In summary, assuming the result of Theorem 1 holds, the effectiveness of PCA in elucidating the geometry of the latent domain from high-dimensional data is closely related to the statistical properties of the random functions (Xj)1≤j≤p such as distinguishability and stationarity. To further explore this relationship, we conduct a simulation study of how the choice of random functions affects the performance of downstream spectral clustering. Let the latent variables (Zi)1≤i≤500 be i.i.d. random elements of a unit circle that form two clusters as in Figure 1a. Let the random functions (Xj)1≤j≤p be i.i.d. real-valued zero-mean Gaussian processes with covariance function exp(−‖x−x′‖22l2), where the length-scale parameter l controls how fast the correlation between two points decays as they grow further apart in Euclidean distance. Further, let Y~i=(Xj(Zi))1≤j≤p denote the p-dimensional embedding of Zi via the random functions, and Ei=(Eij)1≤j≤p the noise vector, where Eij∼i.i.d.N(0,σ2), so that the observed data is Yi=Y~i+Ei. We investigate how well spherical spectral clustering based on (spherical k-means clustering on the spherically projected PCA embeddings of Yi, namely ζi) recovers the cluster structure in the latent variables across different values of σ and l. (a) Latent variables (Zi)1≤i≤500 placed on the unit circle. They are sampled from a two-peak distribution by combining two von Mises distributions (Mardia & Jupp, 2009, Section 3.5.4), Zi∼i.i.d.0.6⋅VM(0,4)+0.4⋅VM(π,4). The latent variables form two clusters that align with the two peaks of the sampling distribution and are distinguished by colour; (b) histogram of latent variables according to their radians; (c) miscluster rate for various pairs of p and σ, fixed parameters: l=0.5,n=500,s=30; (d) miscluster rate for various pairs of p and l, fixed parameters: σ=1,n=500,s=30. We observe that increasing σ slows down the convergence of (ζi)1≤i≤500 to (ϕ(Zi))1≤i≤500, masking the effect of (Y~i)1≤i≤500 and subsequently degrading the spectral clustering performance (Figure 1c). Furthermore, a smaller l implies weaker correlations among (Y~i)1≤i≤500, making Y~i behave more like pure noise, while a larger l implies stronger correlations among (Y~i)1≤i≤500, causing them to concentrate together and become less distinguishable. As the spectral clustering performance in Figure 1d demonstrates, the latent cluster structure is best preserved for values of l where the covariance function correctly assigns discernibly higher correlation for within-cluster pairs compared to between-cluster pairs. From a methodological perspective, the consistency result of Theorem 1 can be complemented by distributional results in certain cases, allowing for hypothesis testing of the latent domain. As a concrete example, assume Y is the sum of a matrix Y~, which is deterministic conditioned on (Zi)1≤i≤n, and an i.i.d. zero-mean noise matrix E. Namely, this is a special case of the Latent Metric Model (LMM) where the functions (Xj)1≤j≤p are deterministic. Then, a distributional theory similar to that in Chang and Cape (2025) can be derived to test a baseline null hypothesis such as H0: Zis a singleton set, via the distribution of ‖sgn(uYT1n)uY−1n1n‖2,∞, where (s12,uY) is the leading eigenpair of YYT, and 1n is the length-n, all-ones vector associated with the leading left singular vector of Y~ under the proposed null where Y~ has n repeated rows. Extending such singular vector distributional results to the more general LMM is nontrivial, as there are multiple sources of randomness, and ϕ is generally not computable, nor intended to be explicitly computed. Hence, identifying and testing null hypotheses via principal components for general manifolds is a potential direction for future research. The code for simulations are available upon request from the author, Xiaoyu Lei.