The impact of incomplete knowledge on evaluation: an experimental benchmark for protein function prediction

The impact of incomplete knowledge on evaluation: an experimental benchmark for protein function prediction
复制标题

DOI:
10.1093/bioinformatics/btp397
复制
发表时间:
2009-09-15
期刊:
影响因子:
5.8
通讯作者:
Troyanskaya, Olga G.
Troyanskaya, Olga G.
中科院分区:
生物学3区
文献类型:
--
作者:
Huttenhower, Curtis;Hibbs, Matthew A.;Troyanskaya, Olga G.

文献摘要

被引文献

相似文献

动机:迅速扩展的高度信息基因组数据的存储库已经对生物网络的蛋白质功能预测和推断的方法产生了越来越多的兴趣。在这些任务上,成功应用监督机器学习需要蛋白质功能的黄金标准:可信赖的一组正确的示例,可用于通过交叉验证或其他统计方法评估性能。由于基因注释对于即使是最佳研究模型生物也不完整,因此可以提出质疑此类评估的生物学可靠性。分析:我们通过通过对线粒体生物发生的蛋白质函数预测进行构建和分析基于实验的基于实验的金标准来解决这种关注。在酿酒酵母中。具体而言,我们确定(i)当前的机器学习方法能够从不完整的黄金标准中概括和预测新型生物学,并且(ii)功能不完整的功能注释对机器学习性能的评估产生不利影响。尽管面对不完整的数据,计算方法的性能要比预测的要好,但竞争方法的相对比较甚至是使用相同培训数据的人,而稀疏的黄金标准也是如此。不完整的知识导致单个方法的表现被差异地低估,从而导致误导性能评估。我们为酵母线粒体提供了基准的金标准,以补充当前数据库,并对我们的实验结果进行分析,以期在将来的比较评估中减轻这些影响。
Motivation: Rapidly expanding repositories of highly informative genomic data have generated increasing interest in methods for protein function prediction and inference of biological networks. The successful application of supervised machine learning to these tasks requires a gold standard for protein function: a trusted set of correct examples, which can be used to assess performance through cross-validation or other statistical approaches. Since gene annotation is incomplete for even the best studied model organisms, the biological reliability of such evaluations may be called into question.Results: We address this concern by constructing and analyzing an experimentally based gold standard through comprehensive validation of protein function predictions for mitochondrion biogenesis in Saccharomyces cerevisiae. Specifically, we determine that (i) current machine learning approaches are able to generalize and predict novel biology from an incomplete gold standard and (ii) incomplete functional annotations adversely affect the evaluation of machine learning performance. While computational approaches performed better than predicted in the face of incomplete data, relative comparison of competing approaches even those employing the same training data-is problematic with a sparse gold standard. Incomplete knowledge causes individual methods' performances to be differentially underestimated, resulting in misleading performance evaluations. We provide a benchmark gold standard for yeast mitochondria to complement current databases and an analysis of our experimental results in the hopes of mitigating these effects in future comparative evaluations.