Gene name ambiguity of eukaryotic nomenclatures

Gene name ambiguity of eukaryotic nomenclatures
复制标题

DOI:
10.1093/bioinformatics/bth496
复制
发表时间:
2005-01-15
期刊:
影响因子:
5.8
通讯作者:
Friedman, C
Friedman, C
中科院分区:
生物学3区
文献类型:
--
作者:
Chen, LF;Liu, HF;Friedman, C

文献摘要

被引文献

相似文献

动机:随着越来越多的科学文献在网上发表,这些知识的有效管理和重用已成为问题。自然语言处理(NLP)可能是一种潜在的解决方案,可以及时提取、结构化和组织网络文献中的生物医学信息。一项基本任务是识别和识别文本中的基因组实体。 “识别”可以通过模式匹配和机器学习来完成。但对于“识别”来说,这些技术还不够。为了识别基因组实体,NLP 需要一个全面的资源,用于指定和分类文本中出现的基因组实体,并将它们与规范化术语和唯一标识符相关联,以便很好地定义提取的实体。在线生物数据库是创建此类词汇资源的绝佳资源。然而,基因名称歧义是一个严重的问题,因为它影响基因实体的正确识别。在本文中,我们探讨了问题的严重程度并提出了解决方法。结果:我们从 21 种生物体中获取了基因信息,并量化了物种内、物种间、英语单词和医学术语的命名歧义。当保留(字母)大小写时,官方符号显示出可忽略不计的种内歧义(0.02%)以及与一般英语单词(0.57%)和医学术语(1.01%)的适度歧义。相比之下,跨物种模糊度较高(14.20%)。基因同义词的包含大大增加了物种内的歧义性,而全名则极大地导致了基因医学术语的歧义性。然后创建了涵盖 21 种生物体基因信息的综合词汇资源,并通过使用简单的字符串匹配程序处理与小鼠模型生物体相关的 45 000 个摘要,同时忽略也是英文单词的大小写和基因名称来识别基因名称。我们发现,85.1% 正确检索的小鼠基因与其他基因名称不明确。当包含也是英文单词的基因名称时,检索到了 233% 的额外“基因”实例,其中大部分是误报。我们还发现,作者更喜欢在其出版物中使用同义词(74.7%)而不是官方符号(17.7%)或全名(7.6%)。
Motivation: With more and more scientific literature published online, the effective management and reuse of this knowledge has become problematic. Natural language processing (NLP) may be a potential solution by extracting, structuring and organizing biomedical information in online literature in a timely manner. One essential task is to recognize and identify genomic entities in text. 'Recognition' can be accomplished using pattern matching and machine learning. But for 'identification' these techniques are not adequate. In order to identify genomic entities, NLP needs a comprehensive resource that specifies and classifies genomic entities as they occur in text and that associates them with normalized terms and also unique identifiers so that the extracted entities are well defined. Online organism databases are an excellent resource to create such a lexical resource. However, gene name ambiguity is a serious problem because it affects the appropriate identification of gene entities. In this paper, we explore the extent of the problem and suggest ways to address it.Results: We obtained gene information from 21 organisms and quantified naming ambiguities within species, across species, with English words and with medical terms. When the case (of letters) was retained, official symbols displayed negligible intra-species ambiguity (0.02%) and modest ambiguities with general English words (0.57%) and medical terms (1.01%). In contrast, the across-species ambiguity was high (14.20%). The inclusion of gene synonyms increased intra-species ambiguity substantially and full names contributed greatly to gene-medical-term ambiguity. A comprehensive lexical resource that covers gene information for the 21 organisms was then created and used to identify gene names by using a straightforward string matching program to process 45 000 abstracts associated with the mouse model organism while ignoring case and gene names that were also English words. We found that 85.1% of correctly retrieved mouse genes were ambiguous with other gene names. When gene names that were also English words were included, 233% additional 'gene' instances were retrieved, most of which were false positives. We also found that authors prefer to use synonyms (74.7%) to official symbols (17.7%) or full names (7.6%) in their publications.