Non-linear Similarity Learning for Semantic Compositionality

Non-linear Similarity Learning for Semantic Compositionality
复制标题

DOI:
10.1527/tjsai.o-fa2
复制
发表时间:
2016
影响因子:
--
通讯作者:
Masashi Tsubaki;M. Shimbo;Yuji Matsumoto
Masashi Tsubaki;M. Shimbo;Yuji Matsumoto
中科院分区:
--
文献类型:
--
作者:
Masashi Tsubaki;M. Shimbo;Yuji Matsumoto

文献摘要

相似文献

文本数据(例如,单词、短语、句子和文档)之间的语义相似性概念在信息检索、分类和提取等自然语言处理(NLP)应用中起着重要作用。最近,使用分布式和分布式模型的词向量空间变得流行起来。虽然词向量提供了很好的词间相似度度量,但从单个词的组成中得出的短语和句子相似度仍然是一个难题。为了解决这个问题,我们专注于在一个比底层词向量空间具有更高表征能力的空间中表示和学习句子的语义相似性。本文提出了一种新的组合性非线性相似学习方法。该方法通过隐式核函数在高维空间中对句子进行相似性学习来学习单词表示,无需在高维空间中显式计算句子向量,就能以较低的成本获得新的单词表示。此外,请注意,我们的方法不同于深度学习,如递归神经网络(rnn)和长短期记忆(LSTM)。我们的目标是设计一种词表示学习,将低维空间(即神经网络)中的嵌入句子结构与高维空间(即核方法)中句子语义的非线性相似性学习相结合。在预测两句语义相似度的任务上(SemEval 2014, task 1),我们的方法优于线性基线、特征工程方法和rnn,并与各种LSTM模型取得了竞争结果。
The notion of semantic similarity between text data (e.g., words, phrases, sentences, and documents) plays an important role in natural language processing (NLP) applications such as information retrieval, classification, and extraction. Recently, word vector spaces using distributional and distributed models have become popular. Although word vectors provide good similarity measures between words, phrasal and sentential similarities derived from composition of individual words remain as a difficult problem. To solve the problem, we focus on representing and learning the semantic similarity of sentences in a space that has a higher representational power than the underlying word vector space. In this paper, we propose a new method of non-linear similarity learning for compositionality. With this method, word representations are learned through the similarity learning of sentences in a high-dimensional space with implicit kernel functions, and we can obtain new word representations inexpensively without explicit computation of sentence vectors in the high-dimensional space. In addition, note that our approach differs from that of deep learning such as recursive neural networks (RNNs) and long short-term memory (LSTM). Our aim is to design a word representation learning which combines the embedding sentence structures in a low-dimensional space (i.e., neural networks) with non-linear similarity learning for the sentence semantics in a high-dimensional space (i.e., kernel methods). On the task of predicting the semantic similarity of two sentences (SemEval 2014, task 1), our method outperforms linear baselines, feature engineering approaches, RNNs, and achieve competitive results with various LSTM models.