Automatic resolution of ambiguous terms based on machine learning and conceptual relations in the UMLS

Automatic resolution of ambiguous terms based on machine learning and conceptual relations in the UMLS
复制标题

DOI:
10.1197/jamia.m1101
复制
发表时间:
2002-11-01
影响因子:
6.4
通讯作者:
Friedman, C
Friedman, C
中科院分区:
管理学2区
文献类型:
--
作者:
Liu, HF;Johnson, SB;Friedman, C

文献摘要

被引文献

相似文献

动力。UMLS已被用于自然语言处理应用,例如信息检索和信息提取系统。自由文本到UMLS概念的映射对于这些应用程序很重要。为了改进映射,我们需要一种方法来消除包含多个UMLS概念的术语的歧义。在普通英语领域,机器学习技术已经被应用于语义标注语料库,在这些语料库中,歧义术语的意义(或概念)已经被标注(主要是人工标注)。然后,语义消歧分类器被导出以自动确定那些歧义术语的含义(或概念)。然而,手动标注语料库是一项昂贵的任务。本文提出了一种基于MEDLINE摘要的语义标注语料库的自动构建方法。对于表示多个UMLS概念的术语W,提取包含W的MEDLINE摘要的集合。对于集合中的每个摘要,自动识别在UMLS中定义的与W有关系的概念的出现。然后,基于这些识别的概念,推导出标注了Ware的意义的语义标注语料库。该方法在一组35个频繁出现的模棱两可的生物医学缩写上进行了评估,使用的是自动派生的黄金标准集。使用查准率和召回率来衡量语义标注语料库的质量。派生语义标注语料库的总体准确率为92.9%,总体召回率为47.4%。去掉稀有义项,忽略义项相近的缩略语,总体准确率为96.8%,总召回率为50.6%。在将自由文本映射到UMLS概念时,可以使用UMLS概念关系和MEDLINE摘要自动获取消解歧义所需的知识。
Motivation. The UMLS has been used in natural language processing applications such as information retrieval and information extraction systems. The mapping of free-text to UMLS concepts is important for these applications. To improve the mapping, we need a method to disambiguate terms that possess multiple UMLS concepts. In the general English domain, machine-learning techniques have been applied to sense-tagged corpora, in which senses (or concepts) of ambiguous terms have been annotated (mostly manually). Sense disambiguation classifiers are then derived to determine senses (or concepts) of those ambiguous terms automatically However, manual annotation of a corpus is an expensive task. We propose an automatic method that constructs sense-tagged corpora for ambiguous terms in the UMLS using MEDLINE abstracts.Methods. For a term W that represents multiple UMLS concepts, a collection of MEDLINE abstracts that contain W is extracted. For each abstract in the collection, occurrences of concepts that have relations with W as defined in the UMLS are automatically identified. A sense-tagged corpus, in which senses of Ware annotated, is then derived based on those identified concepts. The method was evaluated on a set of 35 frequently occurring ambiguous biomedical abbreviations using a gold standard set that was automatically derived. The quality of the derived sense-tagged corpus was measured using precision and recall.Results. The derived sense-tagged corpus had an overall precision of 92.9% and an overall recall of 47.4%. After removing rare senses and ignoring abbreviations with closely related senses, the overall precision was 96.8% and the overall recall was 50.6%.Conclusions. UMLS conceptual relations and MEDLINE abstracts can be used to automatically acquire knowledge needed for resolving ambiguity when mapping free-text to UMLS concepts.