Modeling gene expression regulatory networks with the sparse vector autoregressive model.

Modeling gene expression regulatory networks with the sparse vector autoregressive model.
复制标题

DOI:
10.1186/1752-0509-1-39
复制
发表时间:
2007-08-30
影响因子:
--
通讯作者:
Ferreira CE
Ferreira CE
中科院分区:
生物2区
文献类型:
--
作者:
Fujita A;Sato JR;Garay-Malpartida HM;Yamaguchi R;Miyano S;Sogayar MC;Ferreira CE

文献摘要

参考文献

被引文献

相似文献

为了理解重要生物过程的分子机制,需要详细描述所涉及的基因产物网络。为了定义和理解这样的分子网络,在文献中提出了一些统计方法来估计基因调控网络的时间序列微阵列数据。然而,仍有几个问题需要克服。首先,除了基因之间的相关性之外,还需要推断信息流。其次,我们通常尝试从大量基因(参数)中识别大型网络,这些基因(参数)来自少量的微阵列实验(样本)。由于这种情况在生物信息学中相当常见,因此很难使用对大型基因-基因网络建模的方法进行统计测试。此外,大多数模型都是基于使用聚类技术的降维,因此,所得到的网络不是基因-基因网络,而是模块-模块网络。在这里,我们提出了稀疏向量自回归模型作为解决这些问题。我们应用稀疏向量自回归模型估计基因调控网络的基础上获得的基因表达谱的时间序列微阵列实验。通过大量的模拟,通过将SVAR方法应用于人工调控网络,我们表明,即使在样本数量小于基因数量的条件下,SVAR也可以推断出真正的正边缘。此外,可以控制假阳性,与文献中描述的基于等级或评分函数的其他方法相比,这是一个显著的优点。通过将SVAR应用于实际HeLa细胞周期基因表达数据,我们能够识别众所周知的转录因子靶点。所提出的SVAR方法是能够在频繁的情况下,其中的样本数量低于基因的数量的基因调控网络建模,使得它可以自然地推断部分格兰杰因果关系,没有任何先验信息。此外,我们提出了一个统计测试,以控制错误的发现率,这是以前不可能使用其他基因调控网络模型。
To understand the molecular mechanisms underlying important biological processes, a detailed description of the gene products networks involved is required. In order to define and understand such molecular networks, some statistical methods are proposed in the literature to estimate gene regulatory networks from time-series microarray data. However, several problems still need to be overcome. Firstly, information flow need to be inferred, in addition to the correlation between genes. Secondly, we usually try to identify large networks from a large number of genes (parameters) originating from a smaller number of microarray experiments (samples). Due to this situation, which is rather frequent in Bioinformatics, it is difficult to perform statistical tests using methods that model large gene-gene networks. In addition, most of the models are based on dimension reduction using clustering techniques, therefore, the resulting network is not a gene-gene network but a module-module network. Here, we present the Sparse Vector Autoregressive model as a solution to these problems. We have applied the Sparse Vector Autoregressive model to estimate gene regulatory networks based on gene expression profiles obtained from time-series microarray experiments. Through extensive simulations, by applying the SVAR method to artificial regulatory networks, we show that SVAR can infer true positive edges even under conditions in which the number of samples is smaller than the number of genes. Moreover, it is possible to control for false positives, a significant advantage when compared to other methods described in the literature, which are based on ranks or score functions. By applying SVAR to actual HeLa cell cycle gene expression data, we were able to identify well known transcription factor targets. The proposed SVAR method is able to model gene regulatory networks in frequent situations in which the number of samples is lower than the number of genes, making it possible to naturally infer partial Granger causalities without any a priori information. In addition, we present a statistical test to control the false discovery rate, which was not previously possible using other gene regulatory network models.
DOI: 10.1038/377646a0
发表时间: 1995-10-19
期刊: NATURE
影响因子: 64.8
作者:
BUCKBINDER, L;TALBOTT, R;KLEY, N
通讯作者: KLEY, N
DOI: 10.1074/jbc.270.52.31129
发表时间: 1995-12-29
影响因子: 4.8
作者:
Brown, RT;Ades, IZ;Nordan, RP
通讯作者: Nordan, RP
DOI: 10.1098/rstb.2005.1641
发表时间: 2005-05-29
影响因子: 6.3
作者:
Eichler, M
通讯作者: Eichler, M
DOI: 10.1109/tnn.1997.641482
发表时间: 1997-01-01
影响因子: --
作者:
Cherkassky, V
通讯作者: Cherkassky, V
DOI: 10.1214/009053604000000256
发表时间: 2004-06-01
影响因子: 4.5
作者:
Fan, JQ;Peng, H
通讯作者: Peng, H