用于嵌套数据分析的贝叶斯树结构双层聚类

Bayesian Tree-Structured Two-Level Clustering for Nested Data Analysis

Journal of Computational and Graphical Statistics · 2022
被引 0
ABS 3

中文导读

提出一种非参数贝叶斯模型,对嵌套数据实现树结构双层聚类,自动学习观测层聚类的潜在层次,无需预设聚类数或树结构,适用于多源图像和多主体单细胞表达数据。

Abstract

Data integration plays a crucial role in the era of big data. The nested data are a combined set of observations from multiple sources and exhibit heterogeneity both at the source level and at the observational level. The complex nature makes it challenging to reasonably visualize and jointly analyze the nested data. In this article, we present a nonparametric Bayesian model to implement the tree-structured two-level clustering for nested data analysis. The two-level clustering is used to tease out the heterogeneity existing in the sources and observations, while a tree-structured prior is employed to model the latent hierarchy for clusters at the observational level. The proposed Bayesian model is flexible as it does not require an exact specification of cluster numbers or tree width/depth, and it can automatically learn the underlying tree structures among clusters of observations, thus, offering an insightful visualization of the nested data. We further provide a rigorous posterior sampling scheme via the partially collapsed Gibbs sampler and show the performance of the proposed method using simulation studies. Finally, the applications to two different types of nested data (multi-source image data and multi-subject single-cell expression data) demonstrate the advantages of the proposed Bayesian method. Supplementary materials for this article are available online.

贝叶斯统计聚类分析数据挖掘机器学习