Integrating structured biological data by Kernel Maximum Mean Discrepancy

Integrating structured biological data by Kernel Maximum Mean Discrepancy
复制标题

DOI:
10.1093/bioinformatics/btl242
复制
发表时间:
2006-07-01
期刊:
影响因子:
5.8
通讯作者:
Smola, Alex J.
Smola, Alex J.
中科院分区:
生物学3区
文献类型:
--
作者:
Borgwardt, Karsten M.;Gretton, Arthur;Smola, Alex J.

文献摘要

被引文献

相似文献

动机:生物信息学中数据整合的许多问题可以归结为一个共同的问题:两组观测结果是否由相同的分布产生?我们提出了一个基于核的统计测试这个问题,基于这样一个事实,即两个分布是不同的,当且仅当存在至少一个函数具有不同的期望上的两个分布。因此,我们使用函数平均值之间的最大差异作为检验统计量的基础。最大平均差异(MMD)可以利用内核技巧,这使得我们不仅可以将其应用于向量,还可以将其应用于字符串,序列,图形和其他常见的结构化数据类型在分子生物学中产生。测试微阵列数据的跨平台可比性,癌症诊断,以及两种不同蛋白质功能分类模式的基于数据内容的模式匹配。在所有这些实验中,包括高维的,MMD是非常准确的,在寻找从相同的分布产生的样本,并优于其最好的competitors.Conclusions:我们已经定义了一个新的统计测试是否两个样本是来自相同的分布,兼容的多变量和结构化数据,这是快速,易于实现,工作良好,我们的实验证实。
Motivation: Many problems in data integration in bioinformatics can be posed as one common question: Are two sets of observations generated by the same distribution? We propose a kernel-based statistical test for this problem, based on the fact that two distributions are different if and only if there exists at least one function having different expectation on the two distributions. Consequently we use the maximum discrepancy between function means as the basis of a test statistic.The Maximum Mean Discrepancy (MMD) can take advantage of the kernel trick, which allows us to apply it not only to vectors, but strings, sequences, graphs, and other common structured data types arising in molecular biology.Results: We study the practical feasibility of an MMD-based test on three central data integration tasks: Testing cross-platform comparability of microarray data, cancer diagnosis, and data-content based schema matching for two different protein function classification schemas. In all of these experiments, including high-dimensional ones, MMD is very accurate in finding samples that were generated from the same distribution, and outperforms its best competitors.Conclusions: We have defined a novel statistical test of whether two samples are from the same distribution, compatible with both multivariate and structured data, that is fast, easy to implement, and works well, as confirmed by our experiments.