GAPSCORE:: finding gene and protein names one word at a time

GAPSCORE:: finding gene and protein names one word at a time
复制标题

DOI:
10.1093/bioinformatics/btg393
复制
发表时间:
2004-01-22
期刊:
影响因子:
5.8
通讯作者:
Altman, RB
Altman, RB
中科院分区:
生物学3区
文献类型:
--
作者:
Chang, JT;Schütze, H;Altman, RB

文献摘要

被引文献

相似文献

动机:新的高通量技术加速了基因和蛋白质知识的积累。然而,许多知识仍然以书面自然语言文本的形式存储。因此,我们开发了一种新的方法,GAPSCORE,用于识别文本中的基因和蛋白质名称。GAPSCORE基于量化其外观、形态和上下文的基因名称统计模型对单词进行评分。结果:GAPSCORE与YAPEX数据集进行比较,部分匹配的F分为82.5%(召回率为83.3%,准确率为81.5%),完全匹配的F分为57.6%(召回率为58.5%,准确率为56.7%)。由于该方法是统计的,用户可以根据需要选择分数临界点来调整性能。
Motivation: New high-throughput technologies have accelerated the accumulation of knowledge about genes and proteins. However, much knowledge is still stored as written natural language text. Therefore, we have developed a new method, GAPSCORE, to identify gene and protein names in text. GAPSCORE scores words based on a statistical model of gene names that quantifies their appearance, morphology and context.Results: We evaluated GAPSCORE against the Yapex data set and achieved an F-score of 82.5% (83.3% recall, 81.5% precision) for partial matches and 57.6% (58.5% recall, 56.7% precision) for exact matches. Since the method is statistical, users can choose score cutoffs that adjust the performance according to their needs.