An unbiased evaluation of gene prioritization tools

An unbiased evaluation of gene prioritization tools
复制标题

DOI:
10.1093/bioinformatics/bts581
复制
发表时间:
2012-12-01
期刊:
影响因子:
5.8
通讯作者:
Moreau, Yves
Moreau, Yves
中科院分区:
生物学3区
文献类型:
--
作者:
Bornigen, Daniela;Tranchevent, Leon-Charles;Moreau, Yves

文献摘要

被引文献

相似文献

动机:基因优先排序的目的是从大量的候选基因中找出最有希望的候选基因,从而最大限度地提高下游验证实验和功能研究的产量和生物学相关性。在过去的几年中,已经定义了几个基因优先级工具,其中一些已经通过免费提供的网络工具实现和提供。在这项研究中,我们的目标是比较八个公开可用的优先级工具对新数据的预测性能。我们已经进行了一项分析,其中42个最近报道的疾病基因协会的文献被用来基准这些工具之前,底层databases.Results:交叉验证的回顾性数据提供的性能估计可能是过于乐观,因为一些数据源被污染的疾病基因协会的知识。我们的方法模仿一个新的发现更密切,从而提供更现实的性能估计。然而,存在显著的差异,依赖于更高级的数据集成方案的工具似乎更强大。
Motivation: Gene prioritization aims at identifying the most promising candidate genes among a large pool of candidates-so as to maximize the yield and biological relevance of further downstream validation experiments and functional studies. During the past few years, several gene prioritization tools have been defined, and some of them have been implemented and made available through freely available web tools. In this study, we aim at comparing the predictive performance of eight publicly available prioritization tools on novel data. We have performed an analysis in which 42 recently reported disease-gene associations from literature are used to benchmark these tools before the underlying databases are updated.Results: Cross-validation on retrospective data provides performance estimate likely to be overoptimistic because some of the data sources are contaminated with knowledge from disease-gene association. Our approach mimics a novel discovery more closely and thus provides more realistic performance estimates. There are, however, marked differences, and tools that rely on more advanced data integration schemes appear more powerful.