Embedding of semantic predications.

Embedding of semantic predications.
复制标题

DOI:
10.1016/j.jbi.2017.03.003
复制
发表时间:
2017-04
影响因子:
4.5
通讯作者:
Widdows D
Widdows D
中科院分区:
医学3区
文献类型:
--
作者:
Cohen T;Widdows D

文献摘要

被引文献

相似文献

本文关注的是从结构化知识中生成生物医学概念的分布式向量表示,其形式为被称为语义谓词的主体-关系-对象三元组。具体来说,我们评估了我们以前为此目的开发的代表性方法(称为基于预测的语义索引(PSI))可能受益于从神经概率语言模型中收集的见解的程度,近年来,神经概率语言模型作为一种从自由文本中生成术语的分布式向量表示的手段,受到了越来越多的欢迎。为此,我们开发了一种新的神经概率方法来编码预测,称为嵌入语义预测(ESP),通过调整Skipgram负采样(SGNS)算法的各个方面来实现这一目的。我们比较ESP和PSI在一些任务,包括恢复编码的信息,语义相似性和相关性的估计,并确定潜在的治疗和有害的关系,使用类比检索和监督学习。我们发现ESP在某些方面具有优势,但不是所有这些任务,揭示了神经概率建模的额外计算工作是合理的。
This paper concerns the generation of distributed vector representations of biomedical concepts from structured knowledge, in the form of subject-relation-object triplets known as semantic predications. Specifically, we evaluate the extent to which a representational approach we have developed for this purpose previously, known as Predication-based Semantic Indexing (PSI), might benefit from insights gleaned from neural-probabilistic language models, which have enjoyed a surge in popularity in recent years as a means to generate distributed vector representations of terms from free text. To do so, we develop a novel neural-probabilistic approach to encoding predications, called Embedding of Semantic Predications (ESP), by adapting aspects of the Skipgram with Negative Sampling (SGNS) algorithm to this purpose. We compare ESP and PSI across a number of tasks including recovery of encoded information, estimation of semantic similarity and relatedness, and identification of potentially therapeutic and harmful relationships using both analogical retrieval and supervised learning. We find advantages for ESP in some, but not all of these tasks, revealing the contexts in which the additional computational work of neural-probabilistic modeling is justified.