Selecting high-dimensional mixed graphical models using minimal AIC or BIC forests.

Selecting high-dimensional mixed graphical models using minimal AIC or BIC forests.
复制标题

DOI:
10.1186/1471-2105-11-18
复制
发表时间:
2010-01-11
期刊:
影响因子:
3
通讯作者:
Labouriau R
Labouriau R
中科院分区:
生物学4区
文献类型:
--
作者:
Edwards D;de Abreu GC;Labouriau R

文献摘要

参考文献

被引文献

相似文献

Chow和Liu表明,多变量离散分布的最大似然树可以使用最大权重生成树算法找到,例如Kruskal算法。该算法的效率使其易于处理高维问题。我们扩展周和刘的方法在两个方面:第一,找到森林优化惩罚似然标准,例如AIC或BIC,第二,处理离散和高斯变量的数据。我们将该方法应用于三个数据集:两个来自基因表达研究,第三个来自基因表达研究的遗传学。最小BIC森林通过提供差异表达基因的试验性网络来补充差异表达的常规分析。在基因表达背景的遗传学中,该方法鉴定了近似DNA标记和基因表达水平的联合分布的网络。该方法通常是有用的,作为理解高维离散和/或连续数据的整体依赖结构的初步步骤。树木和森林是不切实际的生物系统的简单模型,但可以提供有用的见解。用途包括:识别不同的连接组件,可以单独分析(降维);识别邻域以进行更详细的分析;作为具有更大搜索空间的搜索算法的初始模型,例如可分解模型或贝叶斯网络;以及识别感兴趣的特征,例如枢纽节点。
Chow and Liu showed that the maximum likelihood tree for multivariate discrete distributions may be found using a maximum weight spanning tree algorithm, for example Kruskal's algorithm. The efficiency of the algorithm makes it tractable for high-dimensional problems. We extend Chow and Liu's approach in two ways: first, to find the forest optimizing a penalized likelihood criterion, for example AIC or BIC, and second, to handle data with both discrete and Gaussian variables. We apply the approach to three datasets: two from gene expression studies and the third from a genetics of gene expression study. The minimal BIC forest supplements a conventional analysis of differential expression by providing a tentative network for the differentially expressed genes. In the genetics of gene expression context the method identifies a network approximating the joint distribution of the DNA markers and the gene expression levels. The approach is generally useful as a preliminary step towards understanding the overall dependence structure of high-dimensional discrete and/or continuous data. Trees and forests are unrealistically simple models for biological systems, but can provide useful insights. Uses include the following: identification of distinct connected components, which can be analysed separately (dimension reduction); identification of neighbourhoods for more detailed analyses; as initial models for search algorithms with a larger search space, for example decomposable models or Bayesian networks; and identification of interesting features, such as hub nodes.
DOI: 10.1126/science.298.5594.824
发表时间: 2002-10-25
期刊: SCIENCE
影响因子: 56.9
作者:
Milo, R;Shen-Orr, S;Alon, U
通讯作者: Alon, U
DOI: 10.1073/pnas.0807227105
发表时间: 2008-12-09
影响因子: 11.1
作者:
Cho, Byung-Kwan;Barrett, Christian L.;Palsson, Bernhard O.
通讯作者: Palsson, Bernhard O.
DOI: 10.1109/tac.1974.1100705
发表时间: 1974-01-01
影响因子: 6.8
作者:
AKAIKE, H
通讯作者: AKAIKE, H
DOI: 10.1089/cmb.2008.08tt
发表时间: 2009-02-01
影响因子: 1.7
作者:
Castelo, Robert;Roverato, Alberto
通讯作者: Roverato, Alberto
DOI: 10.1186/1471-2105-7-s1-s7
发表时间: 2006-03-20
期刊: BMC bioinformatics
影响因子: 3
作者:
Margolin AA;Nemenman I;Basso K;Wiggins C;Stolovitzky G;Dalla Favera R;Califano A
通讯作者: Califano A