Exploring species-based strategies for gene normalization.

Exploring species-based strategies for gene normalization.
复制标题

DOI:
10.1109/tcbb.2010.48
复制
发表时间:
2010-07
期刊:
IEEE/ACM transactions on computational biology and bioinformatics
影响因子:
--
通讯作者:
Hunter LE
Hunter LE
中科院分区:
其他
文献类型:
--
作者:
Verspoor K;Roeder C;Johnson HL;Cohen KB;Baumgartner WA Jr;Hunter LE

文献摘要

相似文献

我们介绍了一个系统开发的BioCreativeII.5社区评估的信息提取的蛋白质和蛋白质相互作用。本文主要集中在基因规范化的任务,识别蛋白质提到的文本和映射到适当的数据库标识符的基础上上下文线索。我们概述了一个“模糊”的字典查找方法,蛋白质提到检测,正则化的文本匹配到类似的正则化的字典条目。我们描述了几种不同的基因标准化策略,专注于文本中提到的物种或生物体,在整个文档中全局性地和在蛋白质提及的附近局部性地,并提出了一系列系统变化的实验结果,这些系统变化探索了各种标准化策略的有效性,以及外部知识来源的作用。虽然我们的系统在评估中既不是表现最好的系统,也不是表现最差的系统,但基因归一化策略显示出希望,并且该系统提供了探索影响BCII.5任务表现的一些变量的机会。
We introduce a system developed for the BioCreativeII.5 community evaluation of information extraction of proteins and protein interactions. The paper focuses primarily on the gene normalization task of recognizing protein mentions in text and mapping them to the appropriate database identifiers based on contextual clues. We outline a “fuzzy” dictionary lookup approach to protein mention detection that matches regularized text to similarly-regularized dictionary entries. We describe several different strategies for gene normalization that focus on species or organism mentions in the text, both globally throughout the document and locally in the immediate vicinity of a protein mention, and present the results of experimentation with a series of system variations that explore the effectiveness of the various normalization strategies, as well as the role of external knowledge sources. While our system was neither the best nor the worst performing system in the evaluation, the gene normalization strategies show promise and the system affords the opportunity to explore some of the variables affecting performance on the BCII.5 tasks.