Learning causal networks from systems biology time course data: an effective model selection procedure for the vector autoregressive process.

Learning causal networks from systems biology time course data: an effective model selection procedure for the vector autoregressive process.
复制标题

DOI:
10.1186/1471-2105-8-s2-s3
复制
发表时间:
2007-05-03
期刊:
影响因子:
3
通讯作者:
Strimmer K
Strimmer K
中科院分区:
生物学4区
文献类型:
--
作者:
Opgen-Rhein R;Strimmer K

文献摘要

被引文献

相似文献

基于向量自回归(VAR)过程的因果网络是一种很有前途的统计工具,可以用来模拟细胞内的调控相互作用。然而,由于基因组数据的低样本量和高维度,学习这些网络是具有挑战性的。我们提出了一种新的、高效的VAR网络估计方法。这分两个步骤进行:(I)使用分析收缩方法改进VAR回归系数的估计,以及(Ii)通过检验相关的部分相关性来随后的模型选择。在模拟实验中,在小样本量情况下,该方法在真实发现率(正确识别的边缘数量相对于有效边缘的数量)方面优于所有其他考虑的方法。此外,对拟南芥表达时间序列数据的分析得出了一个生物学上敏感的网络。即使在基因组学和蛋白质组学中普遍存在的困难数据情况下,所提出的方法也可以有效地完成大规模VAR因果模型的统计学习。该方法是以R代码实现的,该代码可根据作者的要求获得。
Causal networks based on the vector autoregressive (VAR) process are a promising statistical tool for modeling regulatory interactions in a cell. However, learning these networks is challenging due to the low sample size and high dimensionality of genomic data. We present a novel and highly efficient approach to estimate a VAR network. This proceeds in two steps: (i) improved estimation of VAR regression coefficients using an analytic shrinkage approach, and (ii) subsequent model selection by testing the associated partial correlations. In simulations this approach outperformed for small sample size all other considered approaches in terms of true discovery rate (number of correctly identified edges relative to the significant edges). Moreover, the analysis of expression time series data from Arabidopsis thaliana resulted in a biologically sensible network. Statistical learning of large-scale VAR causal models can be done efficiently by the proposed procedure, even in the difficult data situations prevalent in genomics and proteomics. The method is implemented in R code that is available from the authors on request.