聚类数目的近似置信区间

Approximate Confidence Intervals for the Number of Clusters

Journal of the American Statistical Association · 1989
被引 7
ABS 4

中文导读

针对数据降维目的下的聚类问题,提出基于自助法的近似置信区间来估计最优聚类数目,并通过模拟和实例验证其有效性。

Abstract

Abstract We consider clustering for the purpose of data reduction. Similar objects are grouped together in clusters so that one can then work with the few cluster descriptors instead of the many data points. The quality of any given clustering is measured by a loss function that takes into account both the parsimony of the clustering and the loss of information due to clustering. An optimal clustering can be obtained by minimizing the theoretical loss function. It is shown that a sample version of the loss function and optimal clustering converge strongly to their theoretical counterparts as the sample size tends to infinity. We then develop a bootstrap-based procedure for obtaining approximate confidence bounds on the number of clusters in the “best” clustering. The effectiveness of this procedure is evaluated in a simulation study. An application is presented.

聚类分析数据降维统计推断自助法