Estimating the Number of Clusters in a Data Set Via the Gap Statistic
提出一种称为间隙统计量的方法,用于估计数据集中聚类的数量,通过比较实际聚类内离散度与参考零分布下的预期值,模拟研究表明其通常优于现有方法。
Summary We propose a method (the ‘gap statistic’) for estimating the number of clusters (groups) in a set of data. The technique uses the output of any clustering algorithm (e.g. K-means or hierarchical), comparing the change in within-cluster dispersion with that expected under an appropriate reference null distribution. Some theory is developed for the proposal and a simulation study shows that the gap statistic usually outperforms other methods that have been proposed in the literature.