Deep Sentence Embedding Using Long Short-Term Memory Networks: Analysis and Application to Information Retrieval

Deep Sentence Embedding Using Long Short-Term Memory Networks: Analysis and Application to Information Retrieval
复制标题

DOI:
10.1109/taslp.2016.2520371
复制
发表时间:
2016-04-01
影响因子:
5.4
通讯作者:
Ward, Rabab
Ward, Rabab
中科院分区:
计算机科学2区
文献类型:
--
作者:
Palangi, Hamid;Deng, Li;Ward, Rabab

文献摘要

被引文献

相似文献

针对当前自然语言处理研究中的热点问题--句子嵌入,提出了一种基于长短期记忆(LSTM)神经元的递归神经网络(RNN)模型。LSTM-RNN模型按顺序提取句子中的每个单词,提取其信息,并将其嵌入到语义向量中。由于其捕获长期记忆的能力,LSTM-RNN在遍历句子的过程中积累了越来越丰富的信息,当它到达最后一个词时,网络的隐含层提供了整个句子的语义表示。针对商业搜索引擎记录的用户点击数据,对LSTM-RNN进行弱监督训练。为了了解嵌入过程是如何工作的,需要进行可视化和分析。发现该模型能够自动衰减不重要的词,并检测句子中的显著关键词。此外,发现这些检测到的关键字自动激活LSTM-RNN的不同细胞,其中属于相似主题的词激活相同的细胞。作为句子的语义表示,嵌入向量可用于多种不同的应用。由LSTM-RNN实现的这些自动关键字检测和主题分配能力允许网络执行文档检索,这是一项困难的语言处理任务,其中查询和文档之间的相似性可以通过由LSTM-RNN计算的它们对应的句子嵌入向量之间的距离来测量。在网络搜索任务上,LSTM-RNN嵌入被证明显著优于现有的几种最先进的方法。我们强调,该模型生成的句子嵌入向量对Web文档检索任务特别有用。并与一种常见的句子嵌入方法--段落向量法进行了比较。实验结果表明,在Web文档检索任务中,本文提出的方法明显优于段落向量法。
This paper develops a model that addresses sentence embedding, a hot topic in current natural language processing research, using recurrent neural networks (RNN) with Long Short-Term Memory (LSTM) cells. The proposed LSTM-RNN model sequentially takes each word in a sentence, extracts its information, and embeds it into a semantic vector. Due to its ability to capture long term memory, the LSTM-RNN accumulates increasingly richer information as it goes through the sentence, and when it reaches the last word, the hidden layer of the network provides a semantic representation of the whole sentence. In this paper, the LSTM-RNN is trained in a weakly supervised manner on user click-through data logged by a commercial web search engine. Visualization and analysis are performed to understand how the embedding process works. The model is found to automatically attenuate the unimportant words and detect the salient keywords in the sentence. Furthermore, these detected keywords are found to automatically activate different cells of the LSTM-RNN, where words belonging to a similar topic activate the same cell. As a semantic representation of the sentence, the embedding vector can be used in many different applications. These automatic keyword detection and topic allocation abilities enabled by the LSTM-RNN allow the network to perform document retrieval, a difficult language processing task, where the similarity between the query and documents can be measured by the distance between their corresponding sentence embedding vectors computed by the LSTM-RNN. On a web search task, the LSTM-RNN embedding is shown to significantly outperform several existing state of the art methods. We emphasize that the proposed model generates sentence embedding vectors that are specially useful for web document retrieval tasks. A comparison with a well known general sentence embedding method, the Paragraph Vector, is performed. The results show that the proposed method in this paper significantly outperforms Paragraph Vector method for web document retrieval task.