Gracob: a novel graph-based constant-column biclustering method for mining growth phenotype data.

Gracob: a novel graph-based constant-column biclustering method for mining growth phenotype data.
复制标题

DOI:
10.1093/bioinformatics/btx199
复制
发表时间:
2017-08-15
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Gao X
Gao X
中科院分区:
其他
文献类型:
--
作者:
Alzahrani M;Kuwahara H;Wang W;Gao X

文献摘要

参考文献

相似文献

全基因组基因缺失菌株在胁迫条件下的生长表型分析可以清楚地表明基因的重要性取决于环境条件。从这些高通量数据中系统地鉴定在各种环境条件下具有相似的条件必要性和可分配性模式的基因组,可以阐明生长表型的遗传相互作用是如何响应环境而调节的。我们首先证明,检测这样的“共拟合”基因组可以被铸造作为一个不太好研究的问题,在biclustering,即恒定列biclustering。尽管双聚类技术取得了重大进展,但很少有设计用于挖掘生长表型数据。在这里,我们提出了Gracob,一种新的,高效的基于图的方法,铸造和解决恒定列双聚类问题作为一个最大的集团发现问题的多部图。我们将Gracob与大量广泛使用的双聚类方法进行了比较,这些方法涵盖了旨在检测不同类型双聚类的不同类型算法。Gracob在寻找具有广泛设置的各种合成数据集和E.大肠杆菌、变形菌和酵母菌。我们的程序可在http://sfb.kaust.edu.sa/Pages/Software.aspx免费下载。 补充数据可在Bioinformatics在线获得。
Growth phenotype profiling of genome-wide gene-deletion strains over stress conditions can offer a clear picture that the essentiality of genes depends on environmental conditions. Systematically identifying groups of genes from such high-throughput data that share similar patterns of conditional essentiality and dispensability under various environmental conditions can elucidate how genetic interactions of the growth phenotype are regulated in response to the environment. We first demonstrate that detecting such ‘co-fit’ gene groups can be cast as a less well-studied problem in biclustering, i.e. constant-column biclustering. Despite significant advances in biclustering techniques, very few were designed for mining in growth phenotype data. Here, we propose Gracob, a novel, efficient graph-based method that casts and solves the constant-column biclustering problem as a maximal clique finding problem in a multipartite graph. We compared Gracob with a large collection of widely used biclustering methods that cover different types of algorithms designed to detect different types of biclusters. Gracob showed superior performance on finding co-fit genes over all the existing methods on both a variety of synthetic data sets with a wide range of settings, and three real growth phenotype datasets for E. coli, proteobacteria and yeast. Our program is freely available for download at http://sfb.kaust.edu.sa/Pages/Software.aspx. Supplementary data are available at Bioinformatics online.
DOI: 10.1126/science.1150021
发表时间: 2008-04-18
期刊: SCIENCE
影响因子: 56.9
作者:
Hillenmeyer, Maureen E.;Fung, Eula;Giaever, Guri
通讯作者: Giaever, Guri
DOI: 10.1021/bi7014629
发表时间: 2007-11-06
期刊: BIOCHEMISTRY
影响因子: 2.9
作者:
Kim, Juhan;Copley, Shelley D.
通讯作者: Copley, Shelley D.
DOI: 10.1089/10665270360688075
发表时间: 2003-01-01
影响因子: 1.7
作者:
Ben-Dor, A;Chor, B;Yakhini, Z
通讯作者: Yakhini, Z
DOI: 10.1371/journal.pgen.1002385
发表时间: 2011-11
期刊: PLoS genetics
影响因子: 4.5
作者:
Deutschbauer A;Price MN;Wetmore KM;Shao W;Baumohl JK;Xu Z;Nguyen M;Tamse R;Davis RW;Arkin AP
通讯作者: Arkin AP
大肠杆菌K-12的构造框架,单基因敲除突变体:Keio Collection。
DOI: 10.1038/msb4100050
发表时间: 2006
影响因子: 9.9
作者:
通讯作者: --