稀疏K均值元分析框架用于识别多个转录组研究中的疾病亚型

Meta-Analytic Framework for Sparse K-Means to Identify Disease Subtypes in Multiple Transcriptomic Studies

Journal of the American Statistical Association · 2015
被引 39
ABS 4

中文导读

提出一种稀疏K均值的元分析框架,整合多个转录组研究的数据,通过Lasso正则化和模式匹配函数识别疾病亚型,在白血病和乳腺癌数据中比单研究分析更准确稳定。

Abstract

Disease phenotyping by omics data has become a popular approach that potentially can lead to better personalized treatment. Identifying disease subtypes via unsupervised machine learning is the first step toward this goal. In this article, we extend a sparse K-means method toward a meta-analytic framework to identify novel disease subtypes when expression profiles of multiple cohorts are available. The lasso regularization and meta-analysis identify a unique set of gene features for subtype characterization. An additional pattern matching reward function guarantees consistent subtype signatures across studies. The method was evaluated by simulations and leukemia and breast cancer datasets. The identified disease subtypes from meta-analysis were characterized with improved accuracy and stability compared to single study analysis. The breast cancer model was applied to an independent METABRIC dataset and generated improved survival difference between subtypes. These results provide a basis for diagnosis and development of targeted treatments for disease subgroups. Supplementary materials for this article are available online.

生物信息学机器学习癌症亚型元分析转录组学