Bayesian bi-clustering methods with applications in computational biology

Bayesian bi-clustering methods with applications in computational biology
复制标题

DOI:
10.1214/22-aoas1622
复制
发表时间:
2020-07
期刊:
The Annals of Applied Statistics
影响因子:
--
通讯作者:
Han Yan;Jiexing Wu;Y. Li;Jun S. Liu
Han Yan;Jiexing Wu;Y. Li;Jun S. Liu
中科院分区:
其他
文献类型:
--
作者:
Han Yan;Jiexing Wu;Y. Li;Jun S. Liu

文献摘要

相似文献

当观测数据来自不同的群体并且具有大量特征时,双向聚类是分析生物数据的一种有用的方法。概述了处理高维二聚类问题的一般贝叶斯方法,并提出了三种基于分类数据的贝叶斯二聚类模型,这些模型在建模特征在双聚类上的分布方面增加了复杂性。我们提出的方法适用于广泛的场景:从仅在一小部分特征中区分数据但被大量噪声掩盖的情况,到通过不同的特征集识别不同的数据组的情况,再到数据表现出分层结构的情况。通过仿真研究,我们的方法在识别聚类和恢复特征分布模式方面都优于现有的(双)聚类方法。我们将我们的方法应用于两个基因数据集,尽管我们的方法的应用领域更广。我们的方法在实际数据分析中表现出令人满意的性能,并揭示了聚类级别的关系。
Bi-clustering is a useful approach in analyzing biology data when observations come from heterogeneous groups and have a large number of features. We outline a general Bayesian approach in tackling bi-clustering problems in high dimensions, and propose three Bayesian bi-clustering models on categorical data, which increase in complexities in terms of modeling the distributions of features across bi-clusters. Our proposed methods apply to a wide range of scenarios: from situations where data are distinguished only among a small subset of features but masked by a large amount of noise, to situations where different groups of data are identified by different sets of features, to situations where data exhibits hierarchical structures. Through simulation studies, we show that our methods outperform existing (bi-)clustering methods in both identifying clusters and recovering feature distributional patterns across bi-clusters. We apply our methods to two genetic datasets, though the area of application of our methods is even broader. Our methods show satisfactory performance in real data analysis, and reveal cluster-level relationships.