Gene name identification and normalization using a model organism database

Gene name identification and normalization using a model organism database
复制标题

DOI:
10.1016/j.jbi.2004.08.010
复制
发表时间:
2004-12-01
影响因子:
4.5
通讯作者:
Colombe, JB
Colombe, JB
中科院分区:
医学3区
文献类型:
--
作者:
Morgan, AA;Hirschman, L;Colombe, JB

文献摘要

被引文献

相似文献

生物学现已成为一门信息科学,研究人员越来越依赖专家策划的生物数据库来组织已发表文献的研究结果。我们在此报告一系列与自然语言处理应用相关的实验,以帮助 FlyBase 的管理过程。我们重点列出了一篇文章中讨论的基因和基因产物的标准化形式。我们将其分为两个步骤:在文本中标记基因提及,然后对基因名称进行标准化。对于基因提及标签,我们采用了统计方法。为了提供训练数据,我们能够对相关文章和摘要中的基因列表进行逆向工程,以生成(不完美地)标记有基因提及的文本。然后,我们评估了噪声训练数据的质量(精度为 78%,召回率为 88%)以及基于该噪声数据训练的 HMM 标注器输出的质量(精度为 78%,召回率为 71%)。为了生成标准化基因列表,我们探索了两种方法。首先,我们探索了基于同义词列表的简单模式匹配,以获得高召回率/低精度系统(召回率 95%,精度 2%)。使用一系列过滤器,我们能够将精确度提高到 50%,召回率达到 72%(平衡 F 测量值为 0.59)。我们的第二种方法将 HMM 基因提及标记器与各种过滤器相结合,以消除不明确的提及;该方法的 F 测量值为 0.72(精确度 88%,召回率 61%)。这些实验表明,FlyBase提供的词汇资源足够完整,可以在基因列表任务上实现高召回率,并且规范化需要准确的消歧;不同的标记和标准化策略会在召回率和精确度之间进行权衡。 (C) 2004 Elsevier Inc. 保留所有权利。
\Biology has now become an information science, and researchers are increasingly dependent on expert-curated biological databases to organize the findings from the published literature. We report here on a series of experiments related to the application of natural language processing to aid in the curation process for FlyBase. We focused on listing the normalized form of genes and gene products discussed in an article. We broke this into two steps: gene mention tagging in text, followed by normalization of gene names. For gene mention tagging, we adopted a statistical approach. To provide training data, we were able to reverse engineer the gene lists from the associated articles and abstracts, to generate text labeled (imperfectly) with gene mentions. We then evaluated the quality of the noisy training data (precision of 78%, recall 88%) and the quality of the HMM tagger output trained on this noisy data (precision 78%, recall 71%). In order to generate normalized gene lists, we explored two approaches. First, we explored simple pattern matching based on synonym lists to obtain a high recall/low precision system (recall 95%, precision 2%). Using a series of filters, we were able to improve precision to 50% with a recall of 72% (balanced F-measure of 0.59). Our second approach combined the HMM gene mention tagger with various filters to remove ambiguous mentions; this approach achieved an F-measure of 0.72 (precision 88%, recall 61%). These experiments indicate that the lexical resources provided by FlyBase are complete enough to achieve high recall on the gene list task, and that normalization requires accurate disambiguation; different strategies for tagging and normalization trade off recall for precision. (C) 2004 Elsevier Inc. All rights reserved.