A statistical framework for genomic data fusion

A statistical framework for genomic data fusion
复制标题

DOI:
10.1093/bioinformatics/bth294
复制
发表时间:
2004-11-01
期刊:
影响因子:
5.8
通讯作者:
Noble, WS
Noble, WS
中科院分区:
生物学3区
文献类型:
--
作者:
Lanckriet, GRG;De Bie, T;Noble, WS

文献摘要

被引文献

相似文献

动机:在过去的十年中,对基因组学的新的重点突出了一个特殊的挑战:整合不同的意见,基因组所提供的各种类型的实验data.Results:本文介绍了一个计算框架,整合和借鉴全基因组测量的集合推论。每个数据集通过一个核函数表示,该函数定义了实体对(如基因或蛋白质)之间的广义相似性关系。核表示既灵活又高效,可以应用于许多不同类型的数据。此外,从不同类型的数据导出的核函数可以以简单的方式组合。核方法理论的最新进展已经提供了以最小化统计损失函数的方式执行这种组合的有效算法。这些方法利用半定规划技术,以减少寻找优化核组合的问题,凸优化问题。使用酵母全基因组数据集进行的计算实验,包括氨基酸序列,亲水性概况,基因表达数据和已知的蛋白质-蛋白质相互作用,证明了这种方法的实用性。从所有这些数据中训练的统计学习算法识别特定类别的蛋白质-膜蛋白和核糖体蛋白-比任何单一类型的数据训练的相同算法表现得更好。
Motivation: During the past decade, the new focus on genomics has highlighted a particular challenge: to integrate the different views of the genome that are provided by various types of experimental data.Results: This paper describes a computational framework for integrating and drawing inferences from a collection of genome-wide measurements. Each dataset is represented via a kernel function, which defines generalized similarity relationships between pairs of entities, such as genes or proteins. The kernel representation is both flexible and efficient, and can be applied to many different types of data. Furthermore, kernel functions derived from different types of data can be combined in a straightforward fashion. Recent advances in the theory of kernel methods have provided efficient algorithms to perform such combinations in a way that minimizes a statistical loss function. These methods exploit semidefinite programming techniques to reduce the problem of finding optimizing kernel combinations to a convex optimization problem. Computational experiments performed using yeast genome-wide datasets, including amino acid sequences, hydropathy profiles, gene expression data and known protein-protein interactions, demonstrate the utility of this approach. A statistical learning algorithm trained from all of these data to recognize particular classes of proteins-membrane proteins and ribosomal proteins-performs significantly better than the same algorithm trained on any single type of data.