Disambiguating ambiguous biomedical terms in biomedical narrative text: An unsupervised method

Disambiguating ambiguous biomedical terms in biomedical narrative text: An unsupervised method
复制标题

DOI:
10.1006/jbin.2001.1023
复制
发表时间:
2001-08-01
影响因子:
4.5
通讯作者:
Friedman, C
Friedman, C
中科院分区:
医学3区
文献类型:
--
作者:
Liu, HF;Lussier, YA;Friedman, C

文献摘要

被引文献

相似文献

随着自然语言处理(NLP)技术在生物医学领域的信息提取和概念索引的日益广泛使用,同时需要一种快速有效地在给定上下文中分配模糊生物医学术语的正确含义的方法。词义消歧在生物医学领域的现状是基于上下文材料使用手工规则。这种方法的缺点是(i)手动生成WSD规则是一项耗时且乏味的任务,(ii)随着时间的推移,规则集的维护变得越来越困难,以及(iii)手工制定的规则通常不完整,并且在由专门词汇表组成的新领域中表现不佳。和不同类型的文本。本文提出了一种两阶段的无监督方法来建立一个WSD分类器的歧义生物医学术语W。第一阶段自动创建一个意义标记语料库的W,和第二阶段推导出一个分类器W使用派生的意义标记语料库作为训练集。进行了一项形成性实验,该实验表明,在派生的带有含义标签的语料库上训练的分类器的总体准确率约为97%,其中每个歧义术语的准确率均超过90%。(C)2001 Elsevier Science(美国)。
With the growing use of Natural Language Processing (NLP) techniques for information extraction and concept indexing in the biomedical domain, a method that quickly and efficiently assigns the correct sense of an ambiguous biomedical term in a given context is needed concurrently. The current status of word sense disambiguation (WSD) in the biomedical domain is that handcrafted rules are used based on contextual material. The disadvantages of this approach are (i) generating WSD rules manually is a time-consuming and tedious task, (ii) maintenance of rule sets becomes increasingly difficult over time, and (iii) handcrafted rules are often incomplete and perform poorly in new domains comprised of specialized vocabularies and different genres of text. This paper presents a two-phase unsupervised method to build a WSD classifier for an ambiguous biomedical term W The first phase automatically creates a sense-tagged corpus for W, and the second phase derives a classifier for W using the derived sense-tagged corpus as a training set. A formative experiment was performed, which demonstrated that classifiers trained on the derived sense-tagged corpora achieved an overall accuracy of about 97%, with greater than 90% accuracy for each individual ambiguous term. (C) 2001 Elsevier Science (USA).