K-CDFs:一种基于累积分布函数的非参数聚类算法

K-CDFs: A Nonparametric Clustering Algorithm via Cumulative Distribution Function

Journal of Computational and Graphical Statistics · 2022
被引 1
ABS 3

中文导读

提出一种基于累积分布函数的非参数聚类算法K-CDFs,适用于单变量和多变量数据,能有效检测线性不可分簇,对重尾数据和数据维度不敏感。

Abstract

We propose a novel partitioning clustering procedure based on the cumulative distribution function (CDF), called K-CDFs. For univariate data, the K-CDFs represent the cluster centers by empirical CDFs and assign each observation to the closest center measured by the Crame´r-von Mises distance. The procedure is nonparametric and does not require assumptions on cluster distributions imposed by mixture models. A projection technique is used to generalize the K-CDFs for univariate data to an arbitrary dimension. The proposed procedure has several appealing properties. It is robust to heavy-tailed data, is not sensitive to the data dimensions, does not require moment conditions on data and can effectively detect linearly nonseparable clusters. To implement the K-CDFs, we propose two kinds of algorithms: a greedy algorithm as the classical Lloyd’s algorithm and a spectral relaxation algorithm. We illustrate the finite sample performance of the proposed algorithms through simulation experiments and empirical analyses of several real datasets. Supplementary files for this article are available online.

聚类分析非参数统计累积分布函数无监督学习