基于l1融合惩罚的凸聚类:B辑统计方法论

Convex clustering via l1 fusion penalization Series B Statistical methodology

Journal of the Royal Statistical Society. Series A: Statistics in Society · 2017
被引 0
ABS 3

中文导读

研究了凸聚类框架的大样本性质,提出一种后处理修正方法用于估计总体聚类数,并在单细胞病毒学数据中验证了检测细胞亚群的效果。

Abstract

We study the large sample behaviour of a convex clustering framework, which minimizes the sample within cluster sum of squares under an l₁ fusion constraint on the cluster centroids. This recently proposed approach has been gaining in popularity; however, its asymptotic properties have remained mostly unknown. Our analysis is based on a novel representation of the sample clustering procedure as a sequence of cluster splits determined by a sequence of maximization problems. We use this representation to provide a simple and intuitive formulation for the population clustering procedure. We then demonstrate that the sample procedure consistently estimates its population analogue and we derive the corresponding rates of convergence. The proof conducts a careful simultaneous analysis of a collection of M‐estimation problems, whose cardinality grows together with the sample size. On the basis of the new perspectives gained from the asymptotic investigation, we propose a key post‐processing modification of the original clustering framework. We show, both theoretically and empirically, that the resulting approach can be successfully used to estimate the number of clusters in the population. Using simulated data, we compare the proposed method with existing number‐of‐clusters and modality assessment approaches and obtain encouraging results. We also demonstrate the applicability of our clustering method to the detection of cellular subpopulations in a single‐cell virology study.

聚类分析统计学习高维数据算法