Dynamic Pooling and Unfolding Recursive Autoencoders for Paraphrase Detection

Dynamic Pooling and Unfolding Recursive Autoencoders for Paraphrase Detection
复制标题

DOI:
--
复制
发表时间:
2011-12
期刊:
--
影响因子:
--
通讯作者:
R. Socher;E. Huang;Jeffrey Pennington;A. Ng;Christopher D. Manning
R. Socher;E. Huang;Jeffrey Pennington;A. Ng;Christopher D. Manning
中科院分区:
其他
文献类型:
--
作者:
R. Socher;E. Huang;Jeffrey Pennington;A. Ng;Christopher D. Manning

文献摘要

被引文献

相似文献

释义检测是检查两个句子并确定它们是否具有相同的意思的任务。为了在这项任务中获得较高的准确性,需要对这两个语句进行全面的语法和语义分析。提出了一种基于递归自编码器(RAE)的释义检测方法。我们的无监督rae基于一种新的展开目标,并学习语法树中短语的特征向量。这些特征用于测量两个句子之间的单词和短语的相似性。由于句子的长度可以是任意的,因此得到的相似度量矩阵的大小是可变的。我们引入了一种新的动态池化层,它从可变大小的矩阵中计算固定大小的表示。然后将池化表示用作分类器的输入。我们的方法在具有挑战性的MSRP释义语料库上优于其他最先进的方法。
Paraphrase detection is the task of examining two sentences and determining whether they have the same meaning. In order to obtain high accuracy on this task, thorough syntactic and semantic analysis of the two statements is needed. We introduce a method for paraphrase detection based on recursive autoencoders (RAE). Our unsupervised RAEs are based on a novel unfolding objective and learn feature vectors for phrases in syntactic trees. These features are used to measure the word- and phrase-wise similarity between two sentences. Since sentences may be of arbitrary length, the resulting matrix of similarity measures is of variable size. We introduce a novel dynamic pooling layer which computes a fixed-sized representation from the variable-sized matrices. The pooled representation is then used as input to a classifier. Our method outperforms other state-of-the-art approaches on the challenging MSRP paraphrase corpus.