A categorization approach to automated ontological function annotation

A categorization approach to automated ontological function annotation
复制标题

DOI:
10.1110/ps.062184006
复制
发表时间:
2006-06-01
期刊:
影响因子:
8
通讯作者:
Joslyn, Cliff
Joslyn, Cliff
中科院分区:
生物学3区
文献类型:
--
作者:
Verspoor, Karin;Cohn, Judith;Joslyn, Cliff

文献摘要

被引文献

相似文献

自动功能预测(AFP)方法越来越多地使用知识发现算法来将有关功能未知的蛋白质的序列、结构、文献和/或途径信息映射到功能本体中,通常是基因本体(GO)的(一部分)。虽然有越来越多的方法在这个范例中,评估这种预测算法的准确性的一般问题还没有得到认真解决。我们提出了第一个应用程序的功能预测从蛋白质序列使用POSet本体分类器(POSOC)产生新的注释通过分析来自注释的蛋白质BLAST邻域的GO节点的集合。然后,我们还提出了层次精度和层次召回作为新的评估指标,用于评估层次本体中的任何预测的准确性,并讨论了测试集的蛋白质序列的结果。我们表明,我们的方法提供了显着改善的层次精度(测量的预测是正确的)时,应用到最近的BLAST邻居的目标蛋白质,相比简单地估算该邻居的注释的目标。此外,当我们的方法被应用到一个更广泛的BLAST邻域,层次精度进一步提高。在所有情况下,这种增加的分层精度性能都是以适度的分层召回率(对所有得到预测的注释的测量)为代价购买的。
Automated function prediction (AFP) methods increasingly use knowledge discovery algorithms to map sequence, structure, literature, and/or pathway information about proteins whose functions are unknown into functional ontologies, typically (a portion of) the Gene Ontology (GO). While there are a growing number of methods within this paradigm, the general problem of assessing the accuracy of such prediction algorithms has not been seriously addressed. We present first an application for function prediction from protein sequences using the POSet Ontology Categorizer (POSOC) to produce new annotations by analyzing collections of GO nodes derived from annotations of protein BLAST neighborhoods. We then also present hierarchical precision and hierarchical recall as new evaluation metrics for assessing the accuracy of any predictions in hierarchical ontologies, and discuss results on a test set of protein sequences. We show that our method provides substantially improved hierarchical precision (measure of predictions made that are correct) when applied to the nearest BLAST neighbors of target proteins, as compared with simply imputing that neighborhood's annotations to the target. Moreover, when our method is applied to a broader BLAST neighborhood, hierarchical precision is enhanced even further. In all cases, such increased hierarchical precision performance is purchased at a modest expense of hierarchical recall (measure of all annotations that get predicted at all).