A cascaded approach to normalising gene mentions in biomedical literature.

A cascaded approach to normalising gene mentions in biomedical literature.
复制标题

DOI:
10.6026/97320630002197
复制
发表时间:
2007-12-30
期刊:
影响因子:
1.9
通讯作者:
Keane JA
Keane JA
中科院分区:
其他
文献类型:
--
作者:
Yang H;Nenadic G;Keane JA

文献摘要

相似文献

将文献中提到的基因和蛋白质名称与参考基因组数据库中的唯一识别符联系起来,是获取和整合生物医学领域知识的重要步骤。然而,由于词汇和术语的变异,以及文献中提到的基因名称的歧义,这仍然是一项具有挑战性的任务。我们提出了一种通用和有效的基于规则的方法来将文献中的基因提及链接到参考基因组数据库,其中首先应用了对数据库中的基因同义词和文本中的基因提及的预处理。该映射方法采用级联方法,通过灵活地表示基因同义词词典和在预处理阶段产生的基因提及,将精确匹配、精确匹配和基于标记的近似匹配相结合。我们还考虑了多基因名称的提及和基因名称中各组成部分的排列。对所建议的方法进行了系统的评估,确定了有利于提高基因名称识别的精确度或召回率的步骤。在BioCreAtIVE2数据集(人类基因名称的识别)上的实验结果表明,我们的方法取得了非常令人鼓舞的结果,F-度量高达81.20%。
Linking gene and protein names mentioned in the literature to unique identifiers in referent genomic databases is an essential step in accessing and integrating knowledge in the biomedical domain. However, it remains a challenging task due to lexical and terminological variation, and ambiguity of gene name mentions in documents. We present a generic and effective rule-based approach to link gene mentions in the literature to referent genomic databases, where pre-processing of both gene synonyms in the databases and gene mentions in text are first applied. The mapping method employs a cascaded approach, which combines exact, exact-like and token-based approximate matching by using flexible representations of a gene synonym dictionary and gene mentions generated during the pre-processing phase. We also consider multi-gene name mentions and permutation of components in gene names. A systematic evaluation of the suggested methods has identified steps that are beneficial for improving either precision or recall in gene name identification. The results of the experiments on the BioCreAtIvE2 data sets (identification of human gene names) demonstrated that our methods achieved highly encouraging results with F-measure of up to 81.20%.