基于约束似然的模型聚类中正常组数的确定

Finding the Number of Normal Groups in Model-Based Clustering via Constrained Likelihoods

Journal of Computational and Graphical Statistics · 2017
被引 55
ABS 3

中文导读

针对模型聚类中组数k难以确定的问题,提出一种新的惩罚似然准则,通过约束散度矩阵特征值比来控制模型复杂度,并给出自动选择最优(k,c)组合的完整流程,附有“car-bike”图辅助决策。

Abstract

Deciding the number of clusters k is one of the most difficult problems in cluster analysis. For this purpose, complexity-penalized likelihood approaches have been introduced in model-based clustering, such as the well-known Bayesian information criterion and integrated complete likelihood criteria. However, the classification/mixture likelihoods considered in these approaches are unbounded without any constraint on the cluster scatter matrices. Constraints also prevent traditional EM and CEM algorithms from being trapped in (spurious) local maxima. Controlling the maximal ratio between the eigenvalues of the scatter matrices to be smaller than a fixed constant c ⩾ 1 is a sensible idea for setting such constraints. A new penalized likelihood criterion which takes into account the higher model complexity that a higher value of c entails is proposed. Based on this criterion, a novel and fully automated procedure, leading to a small ranked list of optimal (k, c) couples is provided. A new plot called “car-bike,” which provides a concise summary of the solutions, is introduced. The performance of the procedure is assessed both in empirical examples and through a simulation study as a function of cluster overlap. Supplementary materials for the article are available online.

聚类分析模型聚类信息准则约束似然组数选择