Jointly learning word embeddings using a corpus and a knowledge base.

Jointly learning word embeddings using a corpus and a knowledge base.
复制标题

DOI:
10.1371/journal.pone.0193094
复制
发表时间:
2018
期刊:
影响因子:
3.7
通讯作者:
Kawarabayashi KI
Kawarabayashi KI
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Alsuhaibani M;Bollegala D;Maehara T;Kawarabayashi KI

文献摘要

参考文献

被引文献

相似文献

纯粹使用文本语料库中分布的信息来表示向量空间中单词含义的方法已被证明在各种文本挖掘和自然语言处理(NLP)任务中非常有价值。然而,这些方法仍然忽视了共现上下文中单词之间有价值的语义关系结构。这些有益的语义关系结构包含在手动创建的知识库(KB)中,例如本体和语义词典,其中单词的含义通过定义这些单词之间存在的各种关系来表示。我们结合语料库和知识库中的知识来学习更好的词嵌入。具体来说,我们提出了一种联合词表示学习方法,该方法使用知识库中的知识,并同时预测语料库上下文中两个词的共现。特别是,我们使用语料库来定义我们的目标函数,并受到来自知识库的关系约束。我们进一步利用语料库共现统计提出了两种新颖的方法,即最近邻扩展(NNE)和对冲最近邻扩展(HNE),它们动态扩展知识库,从而导出更多指导优化过程的约束。我们在广泛的基准任务上的实验结果表明,所提出的方法在统计上显着提高了所学习的词嵌入的准确性。它优于仅使用语料库的基线,并报告了对先前提出的许多方法的改进,这些方法将语料库和知识库合并到语义相似性预测和单词类比检测任务中。
Methods for representing the meaning of words in vector spaces purely using the information distributed in text corpora have proved to be very valuable in various text mining and natural language processing (NLP) tasks. However, these methods still disregard the valuable semantic relational structure between words in co-occurring contexts. These beneficial semantic relational structures are contained in manually-created knowledge bases (KBs) such as ontologies and semantic lexicons, where the meanings of words are represented by defining the various relationships that exist among those words. We combine the knowledge in both a corpus and a KB to learn better word embeddings. Specifically, we propose a joint word representation learning method that uses the knowledge in the KBs, and simultaneously predicts the co-occurrences of two words in a corpus context. In particular, we use the corpus to define our objective function subject to the relational constrains derived from the KB. We further utilise the corpus co-occurrence statistics to propose two novel approaches, Nearest Neighbour Expansion (NNE) and Hedged Nearest Neighbour Expansion (HNE), that dynamically expand the KB and therefore derive more constraints that guide the optimisation process. Our experimental results over a wide-range of benchmark tasks demonstrate that the proposed method statistically significantly improves the accuracy of the word embeddings learnt. It outperforms a corpus-only baseline and reports an improvement of a number of previously proposed methods that incorporate corpora and KBs in both semantic similarity prediction and word analogy detection tasks.
DOI: 10.1145/503104.503110
发表时间: 2002-01-01
影响因子: 5.6
作者:
Finkelstein, L;Gabrilovich, E;Ruppin, E
通讯作者: Ruppin, E
DOI: 10.1093/nar/gkh061
发表时间: 2004-01-01
影响因子: 14.9
作者:
Bodenreider, O
通讯作者: Bodenreider, O
DOI: 10.1017/s1351324915000431
发表时间: 2017-01-01
影响因子: 2.5
作者:
Hakami, H.;Bollegala, D.
通讯作者: Bollegala, D.