Learning short-text semantic similarity with word embeddings and external knowledge sources

Learning short-text semantic similarity with word embeddings and external knowledge sources
复制标题

DOI:
10.1016/j.knosys.2019.07.013
复制
发表时间:
2019-10-15
影响因子:
8.8
通讯作者:
Cambria, Erik
Cambria, Erik
中科院分区:
计算机科学1区
文献类型:
--
作者:
Nguyen, Hien T.;Duong, Phuc H.;Cambria, Erik

文献摘要

被引文献

相似文献

本文提出了一种基于短文本相互依存表示的短文本语义相似度确定方法。该方法将每个短文本表示为两个密集向量:前者使用基于预训练词向量的词间相似度构建,后者使用基于外部知识来源的词间相似度构建。我们还开发了一种预处理算法,该算法将共指命名实体链接在一起,并执行分词以保留短语动词和习语的含义。我们在Microsoft Research释义语料库、STS2015和P4PIN三个流行的数据集上对所提出的方法进行了评估,并在不使用自然语言先验知识(如词性标签或解析树)的情况下,在这三个数据集上获得了最先进的结果,这表明短文本对的相互依存表示对于语义文本相似任务是有效和高效的。(C) 2019 Elsevier B.V.版权所有
We present a novel method based on interdependent representations of short texts for determining their degree of semantic similarity. The method represents each short text as two dense vectors: the former is built using the word-to-word similarity based on pre-trained word vectors, the latter is built using the word-to-word similarity based on external sources of knowledge. We also developed a preprocessing algorithm that chains coreferential named entities together and performs word segmentation to preserve the meaning of phrasal verbs and idioms. We evaluated the proposed method on three popular datasets, namely Microsoft Research Paraphrase Corpus, STS2015 and P4PIN, and obtained state-of-the-art results on all three without using prior knowledge of natural language, e.g., part-of-speech tags or parse tree, which indicates the interdependent representations of short text pairs are effective and efficient for semantic textual similarity tasks. (C) 2019 Elsevier B.V. All rights reserved.