Bayesian bi-clustering methods with applications in computational biology
Bayesian bi-clustering methods with applications in computational biology
复制标题
DOI:
10.1214/22-aoas1622
复制
发表时间:
2020-07
期刊:
影响因子:
--
通讯作者:
Han Yan;Jiexing Wu;Y. Li;Jun S. Liu
中科院分区:
文献类型:
--
作者:
Han Yan;Jiexing Wu;Y. Li;Jun S. Liu
Bi-clustering is a useful approach in analyzing biology data when observations come from heterogeneous groups and have a large number of features. We outline a general Bayesian approach in tackling bi-clustering problems in high dimensions, and propose three Bayesian bi-clustering models on categorical data, which increase in complexities in terms of modeling the distributions of features across bi-clusters. Our proposed methods apply to a wide range of scenarios: from situations where data are distinguished only among a small subset of features but masked by a large amount of noise, to situations where different groups of data are identified by different sets of features, to situations where data exhibits hierarchical structures. Through simulation studies, we show that our methods outperform existing (bi-)clustering methods in both identifying clusters and recovering feature distributional patterns across bi-clusters. We apply our methods to two genetic datasets, though the area of application of our methods is even broader. Our methods show satisfactory performance in real data analysis, and reveal cluster-level relationships.