An evaluation of GO annotation retrieval for BioCreAtIvE and GOA.

An evaluation of GO annotation retrieval for BioCreAtIvE and GOA.
复制标题

DOI:
10.1186/1471-2105-6-s1-s17
复制
发表时间:
2005
期刊:
影响因子:
3
通讯作者:
Apweiler R
Apweiler R
中科院分区:
生物学4区
文献类型:
--
作者:
Camon EB;Barrell DG;Dimmer EC;Lee V;Magrane M;Maslen J;Binns D;Apweiler R

文献摘要

被引文献

相似文献

基因本体注释 (GOA) 数据库旨在为 UniProt 知识库中的蛋白质提供高质量的补充 GO 注释。与许多其他生物数据库一样,GOA 的大部分内容都是通过精心手工整理的文献收集的。然而,随着文献数量和需要表征的蛋白质数量的增加,手动处理能力可能会变得过载。因此,通常使用半自动辅助工具来加快管理过程。传统上,GOA 中的电子技术很大程度上依赖于利用现有资源(例如 InterPro)中的知识。然而,近年来,文本挖掘被誉为一种有助于管理过程的潜在有用工具。为了鼓励此类工具的开发,EBI 的 GOA 团队同意参加 BioCreAtIvE(生物学信息提取系统的关键评估)挑战赛的功能注释任务。 BioCreAtIvE 任务 2 是一项实验,旨在测试使用信息检索和提取自动导出的分类是否可以帮助专家生物学家将 GO 词汇注释到 UniProt 知识库中的蛋白质。 GOA 提供了从文献中提取的 9000 多个手动 GO 注释的训练语料库。对于测试集,我们提供了 200 篇新的《生物化学杂志》文章的语料库,用于用 GO 术语注释 286 种人类蛋白质。专家团队手动评估了 9 个参与组的结果,每个组都提供了突出显示的句子来支持他们的 GO 和蛋白质注释预测。在这里,我们给出了评估的生物学视角,解释了我们如何使用文献注释 GO,并提出了一些建议以提高未来文本检索和提取技术的精度。最后,我们提供了第一个手动 GO 管理的注释者间一致性研究的结果,以及对我们当前的电子 GO 注释策略的评估。 GOA数据库目前从文献中提取GO注释的精度为91%到100%,召回率至少为72%。这为文本挖掘系统创造了一个特别高的门槛,在 BioCreAtIvE 任务 2(GO 注释提取和检索)中,初始结果只有 10% 到 20% 的时间精确预测 GO 术语。在下一个 BioCreAtIvE 挑战中,GO 术语文本挖掘的性能和准确性有望得到改进。与此同时,GOA 已经采用的手动和电子 GO 注释策略将提供高质量的注释。
The Gene Ontology Annotation (GOA) database aims to provide high-quality supplementary GO annotation to proteins in the UniProt Knowledgebase. Like many other biological databases, GOA gathers much of its content from the careful manual curation of literature. However, as both the volume of literature and of proteins requiring characterization increases, the manual processing capability can become overloaded. Consequently, semi-automated aids are often employed to expedite the curation process. Traditionally, electronic techniques in GOA depend largely on exploiting the knowledge in existing resources such as InterPro. However, in recent years, text mining has been hailed as a potentially useful tool to aid the curation process. To encourage the development of such tools, the GOA team at EBI agreed to take part in the functional annotation task of the BioCreAtIvE (Critical Assessment of Information Extraction systems in Biology) challenge. BioCreAtIvE task 2 was an experiment to test if automatically derived classification using information retrieval and extraction could assist expert biologists in the annotation of the GO vocabulary to the proteins in the UniProt Knowledgebase. GOA provided the training corpus of over 9000 manual GO annotations extracted from the literature. For the test set, we provided a corpus of 200 new Journal of Biological Chemistry articles used to annotate 286 human proteins with GO terms. A team of experts manually evaluated the results of 9 participating groups, each of which provided highlighted sentences to support their GO and protein annotation predictions. Here, we give a biological perspective on the evaluation, explain how we annotate GO using literature and offer some suggestions to improve the precision of future text-retrieval and extraction techniques. Finally, we provide the results of the first inter-annotator agreement study for manual GO curation, as well as an assessment of our current electronic GO annotation strategies. The GOA database currently extracts GO annotation from the literature with 91 to 100% precision, and at least 72% recall. This creates a particularly high threshold for text mining systems which in BioCreAtIvE task 2 (GO annotation extraction and retrieval) initial results precisely predicted GO terms only 10 to 20% of the time. Improvements in the performance and accuracy of text mining for GO terms should be expected in the next BioCreAtIvE challenge. In the meantime the manual and electronic GO annotation strategies already employed by GOA will provide high quality annotations.