Association measures for estimating semantic similarity and relatedness between biomedical concepts

Association measures for estimating semantic similarity and relatedness between biomedical concepts
复制标题

DOI:
10.1016/j.artmed.2018.08.006
复制
发表时间:
2019-01-01
影响因子:
7.5
通讯作者:
McInnes, Bridget T.
McInnes, Bridget T.
中科院分区:
工程技术1区
文献类型:
--
作者:
Henry, Sam;McQuilkin, Alex;McInnes, Bridget T.

文献摘要

被引文献

相似文献

关联度量量化了观察到的术语对共同出现的可能性与它们的预测共同出现在一起(如果是偶然的话)。这是基于术语的单独出现频率和它们的共同出现频率。关联分数的一个应用是估计语义相关性,这对于许多自然语言处理应用是至关重要的,例如生物医学和临床文档的聚类以及生物医学术语和本体学的开发。在本文中,我们提出了一种方法,产生的生物医学概念之间的关联分数,以估计语义相关性。我们使用同现统计统一医学语言系统(UMLS)的概念,占词汇的变化在同义词的水平,并介绍了一个概念扩展的过程,利用层次信息从UMLS占词汇的变化在hyperplane水平。在几个标准评估数据集上取得了最先进的结果,并对超参数进行了深入分析。
Association measures quantify the observed likelihood a term pair co-occurs versus their predicted co-occurrence together if by chance. This is based both on the terms' individual occurrence frequencies, and their mutual co-occurrence frequencies. One application of association scores is estimating semantic relatedness, which is critical for many natural language processing applications, such as clustering of biomedical and clinical documents and the development of biomedical terminologies and ontololgies. In this paper we propose a method of generating association scores between biomedical concepts to estimate semantic relatedness. We use co-occurrence statistics between Unified Medical Language System (UMLS) concepts to account for lexical variation at the synonymous level, and introduce a process of concept expansion that exploits hierarchical information from the UMLS to account for lexical variation at the hyponymous level. State of the art results are achieved on several standard evaluation datasets, and an in depth analysis of hyper-parameters is presented.