The impact of incomplete knowledge on the evaluation of protein function prediction: a structured-output learning perspective.

The impact of incomplete knowledge on the evaluation of protein function prediction: a structured-output learning perspective.
复制标题

DOI:
10.1093/bioinformatics/btu472
复制
发表时间:
2014-09-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Radivojac P
Radivojac P
中科院分区:
其他
文献类型:
--
作者:
Jiang Y;Clark WT;Friedberg I;Radivojac P

文献摘要

参考文献

被引文献

相似文献

动机:生物大分子的自动功能注释是生物学概念或基因和基因产物本体论术语的计算分配问题。已经开发了许多方法来使用标准化术语(例如基因本体论(GO))在计算上注释基因。但是,关于可以整合不同分子数据以及对这些方法的无偏评估的准确方法的可能性仍然存在的问题。一个重要的问题是,蛋白质的实验注释不完整。这就提出了有关是否以及在哪些程度上可以可靠地用于培训计算模型并估算其性能准确性的问题。 结果:我们研究了不完整的实验注释对蛋白质功能预测性能评估可靠性的影响。使用结构化输出学习框架,我们提供理论分析并进行模拟,以表征不断增长的实验注释对与不同类型方法相对应的性能估计的正确性和稳定性的影响。然后,我们通过模拟GO期限预测的预测,评估和随后的重新评估(其他实验注释后)来分析实际生物学数据。我们的结果与以前的观察结果一致,即不完整和积累的实验注释有可能显着影响准确性评估。我们发现它们的影响反映了预测算法,性能指标和基础本体论之间的复杂相互作用。但是,使用可用的实验数据并在现实的假设下,我们的结果还表明,当前的大规模评估是有意义的,几乎令人惊讶地可靠。 联系人:predrag@indiana.edu 补充信息:可以在线生物信息学上获得补充数据。
Motivation: The automated functional annotation of biological macromolecules is a problem of computational assignment of biological concepts or ontological terms to genes and gene products. A number of methods have been developed to computationally annotate genes using standardized nomenclature such as Gene Ontology (GO). However, questions remain about the possibility for development of accurate methods that can integrate disparate molecular data as well as about an unbiased evaluation of these methods. One important concern is that experimental annotations of proteins are incomplete. This raises questions as to whether and to what degree currently available data can be reliably used to train computational models and estimate their performance accuracy. Results: We study the effect of incomplete experimental annotations on the reliability of performance evaluation in protein function prediction. Using the structured-output learning framework, we provide theoretical analyses and carry out simulations to characterize the effect of growing experimental annotations on the correctness and stability of performance estimates corresponding to different types of methods. We then analyze real biological data by simulating the prediction, evaluation and subsequent re-evaluation (after additional experimental annotations become available) of GO term predictions. Our results agree with previous observations that incomplete and accumulating experimental annotations have the potential to significantly impact accuracy assessments. We find that their influence reflects a complex interplay between the prediction algorithm, performance metric and underlying ontology. However, using the available experimental data and under realistic assumptions, our results also suggest that current large-scale evaluations are meaningful and almost surprisingly reliable. Contact: predrag@indiana.edu Supplementary information: Supplementary data are available at Bioinformatics online.
DOI: 10.1186/gb-2008-9-s1-s2
发表时间: 2008
期刊: Genome biology
影响因子: 12.3
作者:
Peña-Castillo L;Tasan M;Myers CL;Lee H;Joshi T;Zhang C;Guan Y;Leone M;Pagnani A;Kim WK;Krumpelman C;Tian W;Obozinski G;Qi Y;Mostafavi S;Lin GN;Berriz GF;Gibbons FD;Lanckriet G;Qiu J;Grant C;Barutcuoglu Z;Hill DP;Warde-Farley D;Grouios C;Ray D;Blake JA;Deng M;Jordan MI;Noble WS;Morris Q;Klein-Seetharaman J;Bar-Joseph Z;Chen T;Sun F;Troyanskaya OG;Marcotte EM;Xu D;Hughes TR;Roth FP
通讯作者: Roth FP
DOI: 10.1093/bioinformatics/btt228
发表时间: 2013-07-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Clark WT;Radivojac P
通讯作者: Radivojac P
DOI: 10.1038/msb4100129
发表时间: 2007
影响因子: 9.9
作者:
Sharan, Roded;Ulitsky, Igor;Shamir, Ron
通讯作者: Shamir, Ron
DOI: 10.1093/bioinformatics/btp397
发表时间: 2009-09-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Huttenhower, Curtis;Hibbs, Matthew A.;Troyanskaya, Olga G.
通讯作者: Troyanskaya, Olga G.
DOI: 10.1186/1471-2105-5-178
发表时间: 2004-11-18
期刊: BMC bioinformatics
影响因子: 3
作者:
Martin DM;Berriman M;Barton GJ
通讯作者: Barton GJ