Disambiguation of biomedical text using diverse sources of information.

Disambiguation of biomedical text using diverse sources of information.
复制标题

DOI:
10.1186/1471-2105-9-s11-s7
复制
发表时间:
2008-11-19
期刊:
影响因子:
3
通讯作者:
Martinez D
Martinez D
中科院分区:
生物学4区
文献类型:
--
作者:
Stevenson M;Guo Y;Gaizauskas R;Martinez D

文献摘要

被引文献

相似文献

与其他领域的文本一样,生物医学文档包含一系列具有多种可能含义的术语。这些模糊性对生物医学文本的自动处理形成了重大障碍。以前的方法来解决这个问题,利用各种来源的信息,包括语言特征的上下文中使用的歧义术语和特定领域的资源,如UMLS。我们比较了各种信息来源,包括以前使用过的信息来源和一个新的信息来源:MeSH术语。使用标准测试集(NLM-WSD语料库)进行评估。使用语言特征和MeSH术语的组合获得最佳性能。我们的系统的性能超过了以前发表的结果,使用相同的数据集评估系统。生物医学术语的消歧得益于使用来自各种来源的信息。特别是,MeSH术语已被证明是有用的,如果可用,应该使用。
Like text in other domains, biomedical documents contain a range of terms with more than one possible meaning. These ambiguities form a significant obstacle to the automatic processing of biomedical texts. Previous approaches to resolving this problem have made use of various sources of information including linguistic features of the context in which the ambiguous term is used and domain-specific resources, such as UMLS. We compare various sources of information including ones which have been previously used and a novel one: MeSH terms. Evaluation is carried out using a standard test set (the NLM-WSD corpus). The best performance is obtained using a combination of linguistic features and MeSH terms. Performance of our system exceeds previously published results for systems evaluated using the same data set. Disambiguation of biomedical terms benefits from the use of information from a variety of sources. In particular, MeSH terms have proved to be useful and should be used if available.