Bayesian biclustering of gene expression data.

Bayesian biclustering of gene expression data.
复制标题

DOI:
10.1186/1471-2164-9-s1-s4
复制
发表时间:
2008
期刊:
影响因子:
4.4
通讯作者:
Liu JS
Liu JS
中科院分区:
生物学2区
文献类型:
--
作者:
Gu J;Liu JS

文献摘要

被引文献

相似文献

基因表达数据的双聚类搜索基因表达的局部模式。双簇(或双向簇)定义为一组基因,其表达谱在实验条件/样品的子集内相互相似。虽然已经研究了几种双聚类算法,但很少有基于严格的统计模型的。我们开发了一个贝叶斯双聚类模型(BBC),并实现了吉布斯抽样过程的统计推断。我们表明,贝叶斯双聚类模型可以正确地识别多个集群的基因表达数据。通过对模型和实际数据的仿真,证明了BBC算法在鲁棒性和准确性方面均优于其他方法。我们还证明了模型对于两种归一化方法是稳定的,四分位距归一化和最小四分位距归一化。将BBC算法应用于酵母表达数据,我们观察到我们发现的大多数双簇都得到了重要的生物学证据的支持,例如相应启动子序列中基因功能和转录因子结合位点的富集。BBC算法被证明是一个强大的基于模型的biclustering方法,可以发现在微阵列数据的生物学显着的基因条件集群。BBC模型可以很容易地通过Monte Carlo插补处理缺失数据,并有可能扩展到基因转录网络的综合研究。
Biclustering of gene expression data searches for local patterns of gene expression. A bicluster (or a two-way cluster) is defined as a set of genes whose expression profiles are mutually similar within a subset of experimental conditions/samples. Although several biclustering algorithms have been studied, few are based on rigorous statistical models. We developed a Bayesian biclustering model (BBC), and implemented a Gibbs sampling procedure for its statistical inference. We showed that Bayesian biclustering model can correctly identify multiple clusters of gene expression data. Using simulated data both from the model and with realistic characters, we demonstrated the BBC algorithm outperforms other methods in both robustness and accuracy. We also showed that the model is stable for two normalization methods, the interquartile range normalization and the smallest quartile range normalization. Applying the BBC algorithm to the yeast expression data, we observed that majority of the biclusters we found are supported by significant biological evidences, such as enrichments of gene functions and transcription factor binding sites in the corresponding promoter sequences. The BBC algorithm is shown to be a robust model-based biclustering method that can discover biologically significant gene-condition clusters in microarray data. The BBC model can easily handle missing data via Monte Carlo imputation and has the potential to be extended to integrated study of gene transcription networks.