Evaluating UMLS strings for natural language processing

Evaluating UMLS strings for natural language processing
复制标题

评估自然语言处理的 UMLS 字符串

DOI:
--
复制
发表时间:
2001
期刊:
American Medical Informatics Association Annual Symposium
影响因子:
--
通讯作者:
Allen C. Browne
Allen C. Browne
中科院分区:
--
文献类型:
--
作者:
A. McCray;O. Bodenreider;J. Malley;Allen C. Browne

文献摘要

参考文献

被引文献

相似文献

国家医学图书馆的统一医学语言系统(UMLS)是生物医学领域丰富的知识来源。UMLS用于一系列不同应用的研究和开发,包括自然语言处理(NLP)。在本文中,我们调查了UMLS Metathesaurus中的字符串的性质,并评估了它们在NLP中的有用性。我们首先确定一些属性,这些属性可能允许我们预测在语料库中找到或找不到给定字符串的可能性。我们使用统计模型对我们的语料库进行测试,语料库来自MEDLINE数据库。对于一组属性,该模型正确地预测了77%的不属于语料库的字符串和85%的属于语料库的字符串。对于另一组属性,该模型正确地预测了96%的不属于语料库的字符串和29%的属于语料库的字符串。
The National Library of Medicine's Unified Medical Language System (UMLS) is a rich source of knowledge in the biomedical domain. The UMLS is used for research and development in a range of different applications, including natural language processing (NLP). In this paper we investigate the nature of the strings found in the UMLS Metathesaurus and evaluate them for their usefulness in NLP. We begin by identifying a number of properties that might allow us to predict the likelihood of a given string being found or not found in a corpus. We use a statistical model to test these predictors against our corpus, which is derived from the MEDLINE database. For one set of properties the model correctly predicted 77% of the strings that do not belong to the corpus, and 85% of the strings that do belong to the corpus. For another set of properties the model correctly predicted 96% of the strings that do not belong to the corpus and 29% of the strings that do belong to the corpus.
使用统一医学语言系统的知识表示和索引。
DOI: 10.1142/9789814447331_0047
发表时间: 2000
影响因子: --
作者:
Baclawski,K;Cigna,J;Kokar,MM;Mager,P;Indurkhya,B
通讯作者: Indurkhya,B