非参数装袋聚类方法:从依赖分类数据序列中识别潜在结构

Nonparametric bagging clustering methods to identify latent structures from a sequence of dependent categorical data

Computational Statistics and Data Analysis · 2022
被引 5
ABS 3

中文导读

研究并比较了多种非参数装袋聚类方法,用于从一维时间域上观测的依赖分类数据序列中恢复缓慢变化的潜在信号,发现基于熵的复合方法能稳健提升性能。

Abstract

Nonparametric bagging clustering methods are studied and compared to identify latent structures from a sequence of dependent categorical data observed along a one-dimensional (discrete) time domain. The frequency of the observed categories is assumed to be generated by a (slowly varying) latent signal, according to latent state-specific probability distributions. The bagging clustering methods use random tessellations (partitions) of the time domain and clustering of the category frequencies of the observed data in the tessellation cells to recover the latent signal, within a bagging framework. New and existing ways of generating the tessellations and clustering are discussed and combined into different bagging clustering methods. Edge tessellations and adaptive tessellations are the new proposed ways of forming partitions. Composite methods are also introduced, that are using (automated) decision rules based on entropy measures to choose among the proposed bagging clustering methods. The performance of all the methods is compared in a simulation study. From the simulation study it can be concluded that local and global entropy measures are powerful tools in improving the recovery of the latent signal, both via the adaptive tessellation strategies (local entropy) and in designing composite methods (global entropy). The composite methods are robust and overall improve performance, in particular the composite method using adaptive (edge) tessellations.

聚类分析分类数据非参数统计时间序列分析模式识别