Selective integration of multiple biological data for supervised network inference

Selective integration of multiple biological data for supervised network inference
复制标题

DOI:
10.1093/bioinformatics/bti339
复制
发表时间:
2005-05-15
期刊:
影响因子:
5.8
通讯作者:
Asai, K
Asai, K
中科院分区:
生物学3区
文献类型:
--
作者:
Kato, T;Tsuda, K;Asai, K

文献摘要

被引文献

相似文献

动机:从生物数据推断蛋白质网络是计算生物学的核心问题。大多数网络推理方法,包括贝叶斯网络,都采用无监督的方法,其中网络在开始时是完全未知的,并且必须预测所有的边。最近提出了一个更现实的监督框架,假设网络的很大一部分是已知的。我们提出了一种新的基于核的监督图推理方法,该方法基于多种类型的生物数据集,如基因表达、系统发育谱和氨基酸序列。值得注意的是,我们的方法为每种类型的数据集分配了权重,从而选择了信息丰富的数据集。数据选择有助于降低数据收集成本。例如,当必须为其他生物解决类似的网络推理问题时,我们的算法不需要收集排除的数据集。结果:首先,我们将监督网络推理表述为核矩阵补全问题,其中边的推理归结为核矩阵缺失项的估计。然后,提出了一种期望最大化算法来同时推断核矩阵的缺失项和多个数据集的权重。通过引入权重,我们可以有选择地整合多个数据集,从而排除不相关和有噪声的数据集。我们的方法在两个生物网络中得到了良好的测试:代谢网络和蛋白质相互作用网络。
Motivation: Inferring networks of proteins from biological data is a central issue of computational biology. Most network inference methods, including Bayesian networks, take unsupervised approaches in which the network is totally unknown in the beginning, and all the edges have to be predicted. A more realistic supervised framework, proposed recently, assumes that a substantial part of the network is known. We propose a new kernel-based method for supervised graph inference based on multiple types of biological datasets such as gene expression, phylogenetic profiles and amino acid sequences. Notably, our method assigns a weight to each type of dataset and thereby selects informative ones. Data selection is useful for reducing data collection costs. For example, when a similar network inference problem must be solved for other organisms, the dataset excluded by our algorithm need not be collected.Results: First, we formulate supervised network inference as a kernel matrix completion problem, where the inference of edges boils down to estimation of missing entries of a kernel matrix. Then, an expectation-maximization algorithm is proposed to simultaneously infer the missing entries of the kernel matrix and the weights of multiple datasets. By introducing the weights, we can integrate multiple datasets selectively and thereby exclude irrelevant and noisy datasets. Our approach is favorably tested in two biological networks: a metabolic network and a protein interaction network.