A data integration methodology for systems biology

A data integration methodology for systems biology
复制标题

DOI:
10.1073/pnas.0508647102
复制
发表时间:
2005-11-29
影响因子:
11.1
通讯作者:
Bolouri, H
Bolouri, H
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Hwang, D;Rust, AG;Bolouri, H

文献摘要

被引文献

相似文献

不同的实验技术测量系统的不同方面,并具有不同的深度和广度。高通量测定具有固有的高假阳性率和假阴性率。此外,每种技术都包含不同性质的系统性偏差。这些差异使得从多个数据集重建网络变得困难且容易出错。此外,由于生物技术的快速发展,通常没有精心策划的样本数据集,人们可以从中估计数据集成参数。为了解决这些问题,我们开发了数据集成方法,可以处理统计能力、类型、大小和网络覆盖范围不同的多个数据集,而不需要精心策划的训练数据集。我们的方法是通用的,可以应用于整合任何现有和未来技术的数据。在这里,我们概述了我们的方法,然后通过将它们应用于模拟数据集来展示它们的性能。结果表明,这些方法比传统方法更准确地选择真阳性数据元素。在随附的配套文件中,我们证明了我们的方法对生物数据的适用性。我们已经将我们的方法集成到一个名为POINTILLIST的免费开源软件包中。
Different experimental technologies measure different aspects of a system and to differing depth and breadth. High-throughput assays have inherently high false-positive and false-negative rates. Moreover, each technology includes systematic biases of a different nature. These differences make network reconstruction from multiple data sets difficult and error-prone. Additionally, because of the rapid rate of progress in biotechnology, there is usually no curated exemplar data set from which one might estimate data integration parameters. To address these concerns, we have developed data integration methods that can handle multiple data sets differing in statistical power, type, size, and network coverage without requiring a curated training data set. Our methodology is general in purpose and may be applied to integrate data from any existing and future technologies. Here we outline our methods and then demonstrate their performance by applying them to simulated data sets. The results show that these methods select true-positive data elements much more accurately than classical approaches. In an accompanying companion paper, we demonstrate the applicability of our approach to biological data. We have integrated our methodology into a free open source software package named POINTILLIST.