Evaluating semantic relations in neural word embeddings with biomedical and general domain knowledge bases.

Evaluating semantic relations in neural word embeddings with biomedical and general domain knowledge bases.
复制标题

用生物医学和一般领域知识库评估神经单词嵌入的语义关系。

DOI:
10.1186/s12911-018-0630-x
复制
发表时间:
2018-07-23
影响因子:
3.5
通讯作者:
Bian J
Bian J
中科院分区:
医学3区
文献类型:
--
作者:
Chen Z;He Z;Liu X;Bian J

文献摘要

参考文献

被引文献

相似文献

在过去的几年里,神经词嵌入在文本挖掘中得到了广泛的应用。然而,词嵌入的向量表示在使用它们的下游应用程序中大多充当黑匣子,从而限制了它们的可解释性。尽管词嵌入能够捕获自由文本文档中的语义规律,但尚不清楚词嵌入如何表示不同类型的语义关系以及如何从词嵌入中检索语义相关的术语。为了提高词嵌入的透明度和使用它们的应用程序的可解释性,在本研究中,我们提出了一种使用外部知识库评估词嵌入中语义关系的新方法:维基百科、WordNet 和统一医学语言系统(UMLS)。我们使用维基百科中与健康相关的文章训练多个词嵌入,然后评估它们在类比和语义关系术语检索任务中的表现。我们还通过将健康相关维基百科文章的嵌入与一般维基百科文章的嵌入进行比较来评估评估结果是否取决于文本语料库的领域。关于语义关系的检索,我们能够检索语义。与此同时,两种流行的词嵌入方法Word2vec和GloVe在类比检索任务和语义关系检索任务上都获得了可比较的结果,而基于依存关系的词嵌入在这两个任务中的表现要差得多。我们还发现,使用健康相关维基百科文章训练的词嵌入在健康相关关系检索任务中比使用一般维基百科文章训练的词嵌入获得了更好的性能。从这项研究中可以明显看出,词嵌入可以将具有不同语义关系的术语组合在一起。训练语料库的领域确实对词嵌入表示的语义关系有影响。因此,我们建议使用特定领域的语料库来训练特定领域文本挖掘任务的词嵌入。
In the past few years, neural word embeddings have been widely used in text mining. However, the vector representations of word embeddings mostly act as a black box in downstream applications using them, thereby limiting their interpretability. Even though word embeddings are able to capture semantic regularities in free text documents, it is not clear how different kinds of semantic relations are represented by word embeddings and how semantically-related terms can be retrieved from word embeddings. To improve the transparency of word embeddings and the interpretability of the applications using them, in this study, we propose a novel approach for evaluating the semantic relations in word embeddings using external knowledge bases: Wikipedia, WordNet and Unified Medical Language System (UMLS). We trained multiple word embeddings using health-related articles in Wikipedia and then evaluated their performance in the analogy and semantic relation term retrieval tasks. We also assessed if the evaluation results depend on the domain of the textual corpora by comparing the embeddings of health-related Wikipedia articles with those of general Wikipedia articles. Regarding the retrieval of semantic relations, we were able to retrieve semanti. Meanwhile, the two popular word embedding approaches, Word2vec and GloVe, obtained comparable results on both the analogy retrieval task and the semantic relation retrieval task, while dependency-based word embeddings had much worse performance in both tasks. We also found that the word embeddings trained with health-related Wikipedia articles obtained better performance in the health-related relation retrieval tasks than those trained with general Wikipedia articles. It is evident from this study that word embeddings can group terms with diverse semantic relations together. The domain of the training corpus does have impact on the semantic relations represented by word embeddings. We thus recommend using domain-specific corpus to train word embeddings for domain-specific text mining tasks.
DOI: 10.1016/j.jbi.2017.03.016
发表时间: 2017-05
影响因子: 4.5
作者:
He Z;Chen Z;Oh S;Hou J;Bian J
通讯作者: Bian J
DOI: 10.1145/219717.219748
发表时间: 1995-11-01
影响因子: 22.7
作者:
MILLER, GA
通讯作者: MILLER, GA
DOI: 10.1080/00437956.1954.11659520
发表时间: 1954-08-01
影响因子: 0.6
作者:
Harris, Zellig S.
通讯作者: Harris, Zellig S.
DOI: 10.1109/tvcg.2017.2745141
发表时间: 2018-01-01
影响因子: 5.2
作者:
Liu, Shusen;Bremer, Peer-Timo;Pascucci, Valerio
通讯作者: Pascucci, Valerio
DOI: 10.1055/s-0038-1634945
发表时间: 1993-08-01
影响因子: 1.7
作者:
LINDBERG, DAB;HUMPHREYS, BL;MCCRAY, AT
通讯作者: MCCRAY, AT