MCbiclust: a novel algorithm to discover large-scale functionally related gene sets from massive transcriptomics data collections

MCbiclust: a novel algorithm to discover large-scale functionally related gene sets from massive transcriptomics data collections
复制标题

MCbiclust:一种从大量转录组数据集中发现大规模功能相关基因集的新算法

DOI:
10.1101/075374
复制
发表时间:
2016
期刊:
--
影响因子:
--
通讯作者:
Bentham R
Bentham R
中科院分区:
--
文献类型:
--
作者:
Bentham R

文献摘要

参考文献

被引文献

相似文献

随着最近数据收集规模的爆炸性增长,从基因表达数据中了解基本生物学过程的可能性也在增长。然而,为了开发这一潜力,需要能够发现大型共同调节基因网络的新的分析方法。我们发现,目前的方法限制了他们可以在生物异质数据集合中发现的相关基因集的大小,阻碍了对能量代谢、细胞器生物发生和应激反应等多基因控制的基本细胞过程的识别。在这里,我们描述了一种新的双聚类算法,称为大规模相关双聚类(MCbilust),它从具有最大相关基因表达的大数据集中选择样本和基因,从而允许检查复杂网络的调节。该方法已经使用合成数据进行了评估,并应用于大型细菌和癌细胞数据集。我们表明,到目前为止,通过现有技术难以识别的大型双色体是生物相关的,因此MCbilust在分析转录组数据以识别隐藏在数据中的大规模未知效应方面具有巨大的潜力。识别出的巨大双聚体可用于开发改进的基于转录组学的疾病诊断工具,用于基因表达变化引起的疾病,或用于进一步的网络分析,以了解基因型-表型相关性。
The potential to understand fundamental biological processes from gene expression data has grown in parallel with the recent explosion of the size of data collections. However, to exploit this potential, novel analytical methods are required, capable of discovering large co-regulated gene networks. We found current methods limited in the size of correlated gene sets they could discover within biologically heterogeneous data collections, hampering the identification of multi-gene controlled fundamental cellular processes such as energy metabolism, organelle biogenesis and stress responses. Here we describe a novel biclustering algorithm called Massively Correlated Biclustering (MCbiclust) that selects samples and genes from large datasets with maximal correlated gene expression, allowing regulation of complex networks to be examined. The method has been evaluated using synthetic data and applied to large bacterial and cancer cell datasets. We show that the large biclusters discovered, so far elusive to identification by existing techniques, are biologically relevant and thus MCbiclust has great potential in the analysis of transcriptomics data to identify large-scale unknown effects hidden within the data. The identified massive biclusters can be used to develop improved transcriptomics based diagnosis tools for diseases caused by altered gene expression, or used for further network analysis to understand genotype-phenotype correlations.
DOI: 10.1038/nature14587
发表时间: 2015-08-20
期刊: NATURE
影响因子: 64.8
作者:
Perera, RushikaM.;Stoykova, Svetlana;Nicolay, Brandon N.;Ross, Kenneth N.;Fitamant, Julien;Boukhali, Myriam;Lengrand, Justine;Deshpande, Vikram;Selig, Martin K.;Ferrone, Cristina R.;Settleman, Jeff;Stephanopoulos, Gregory;Dyson, Nicholas J.;Zoncu, Roberto;Ramaswamy, Sridhar;Haas, Wilhelm;Bardeesy, Nabeel
通讯作者: Bardeesy, Nabeel
DOI: 10.1016/j.cmpb.2013.07.025
发表时间: 2013-12-01
影响因子: 6.1
作者:
Flores, Jose L.;Inza, Inaki;Calvo, Borja
通讯作者: Calvo, Borja
关于分配和交通问题(摘要)
DOI: 10.1002/nav.3800040112
发表时间: 1957
期刊: Naval Research Logistics Quarterly
影响因子: --
作者:
J. Munkres
通讯作者: J. Munkres
DOI: 10.1007/978-1-61779-400-1_3
发表时间: 2012
期刊: Methods in molecular biology (Clifton, N.J.)
影响因子: --
作者:
Wilhite, Stephen E;Barrett, Tanya
通讯作者: Barrett, Tanya
DOI: 10.1093/nar/gkr359
发表时间: 2011-07
影响因子: 14.9
作者:
Lan A;Smoly IY;Rapaport G;Lindquist S;Fraenkel E;Yeger-Lotem E
通讯作者: Yeger-Lotem E