Clustering risk in nonparametric hidden Markov and I.I.D. models
分析了隐马尔可夫和独立同分布模型中聚类的贝叶斯风险,发现贝叶斯分类器在聚类中接近最优,并给出了非参数设定下插件分类器聚类超额风险的界。
We conduct an in-depth analysis of the Bayes risk of clustering in the context of hidden Markov and i.i.d. models. In both settings, we identify the situations where this risk is comparable to the Bayes risk of classification and those where its minimizer, the Bayes clusterer, can be derived from the Bayes classifier. While we demonstrate that clustering based on the Bayes classifier does not always match the optimal Bayes clusterer, we show that this difference is primarily theoretical and that the Bayes classifier remains nearly optimal for clustering. A key quantity emerges, capturing the fundamental difficulty of both classification and clustering tasks. Furthermore, by leveraging the identifiability of HMMs, we establish bounds on the clustering excess risk of a plug-in Bayes classifier in the general nonparametric setting, offering theoretical justification for its widespread use in practice. Simulations further illustrate our findings.