Data-poor categorization and passage retrieval for gene ontology annotation in Swiss-Prot.

Data-poor categorization and passage retrieval for gene ontology annotation in Swiss-Prot.
复制标题

DOI:
10.1186/1471-2105-6-s1-s23
复制
发表时间:
2005
期刊:
影响因子:
3
通讯作者:
Ruch P
Ruch P
中科院分区:
生物学4区
文献类型:
--
作者:
Ehrler F;Geissbühler A;Jimeno A;Ruch P

文献摘要

被引文献

相似文献

在BioCreative竞赛的背景下,训练数据非常稀疏,我们研究了两个互补的任务:1)给定一个Swiss-Prot三联体,包含一个蛋白质,一个GO(基因本体论)术语和一个相关文章,提取一个证明GO类别分配合理的短段落; 2)给定一个Swiss-Prot对,包含一个蛋白质和一个相关文章,自动分配一组类别。句子是最基本的检索单位。我们的分类器计算每个句子与Swiss-Prot条目提供的GO类别之间的距离。文本分类器计算每个GO术语与文章文本之间的距离。评估报告的基础上,注释者的判断,建立了竞争,并根据平均平均精度的措施计算使用策划样本的瑞士-Prot。我们的系统实现了最好的召回率和精度组合的段落检索和文本分类的官方评估。然而,文本分类的结果远远低于其他数据贫乏的文本分类实验中的结果。在不到20%的情况下,我们提出的最高术语是相关的,而其他生物医学控制词汇的分类,如医学主题词,我们达到了90%以上的精度。我们还观察到,在我们的实验中使用的评分方法,我们的引擎的检索状态值的基础上,表现出有效的置信度估计能力。从比较的角度来看,我们设计的检索和自然语言处理方法的结合,取得了非常有竞争力的性能。尽管我们的系统与数据无关,但其有效性并不亚于数据密集型方法。这些结果表明,整体策略可以使一大类信息提取任务受益,特别是当训练数据缺失时。然而,从用户的角度来看,结果令人失望。需要进一步的调查,以设计适用的最终用户的文本挖掘工具的生物学家。
In the context of the BioCreative competition, where training data were very sparse, we investigated two complementary tasks: 1) given a Swiss-Prot triplet, containing a protein, a GO (Gene Ontology) term and a relevant article, extraction of a short passage that justifies the GO category assignement; 2) given a Swiss-Prot pair, containing a protein and a relevant article, automatic assignement of a set of categories. Sentence is the basic retrieval unit. Our classifier computes a distance between each sentence and the GO category provided with the Swiss-Prot entry. The Text Categorizer computes a distance between each GO term and the text of the article. Evaluations are reported both based on annotator judgements as established by the competition and based on mean average precision measures computed using a curated sample of Swiss-Prot. Our system achieved the best recall and precision combination both for passage retrieval and text categorization as evaluated by official evaluators. However, text categorization results were far below those in other data-poor text categorization experiments The top proposed term is relevant in less that 20% of cases, while categorization with other biomedical controlled vocabulary, such as the Medical Subject Headings, we achieved more than 90% precision. We also observe that the scoring methods used in our experiments, based on the retrieval status value of our engines, exhibits effective confidence estimation capabilities. From a comparative perspective, the combination of retrieval and natural language processing methods we designed, achieved very competitive performances. Largely data-independent, our systems were no less effective that data-intensive approaches. These results suggests that the overall strategy could benefit a large class of information extraction tasks, especially when training data are missing. However, from a user perspective, results were disappointing. Further investigations are needed to design applicable end-user text mining tools for biologists.