A Neural Conversational Model

A Neural Conversational Model
复制标题

DOI:
--
复制
发表时间:
2015-06
期刊:
ArXiv
影响因子:
--
通讯作者:
O. Vinyals;Quoc V. Le
O. Vinyals;Quoc V. Le
中科院分区:
其他
文献类型:
--
作者:
O. Vinyals;Quoc V. Le

文献摘要

被引文献

相似文献

对话建模是自然语言理解和机器智能中的一项重要任务。尽管先前存在一些方法,但它们往往局限于特定领域(例如预订机票),并且需要手工制定规则。在本文中,我们针对该任务提出了一种简单的方法,该方法使用了最近提出的序列到序列框架。我们的模型通过在对话中根据之前的一个或多个句子预测下一个句子来进行对话。我们模型的优势在于它可以进行端到端的训练,因此所需的手工制定规则要少得多。我们发现,给定一个大型的对话训练数据集,这个简单的模型可以生成简单的对话。我们的初步结果表明,尽管优化的是错误的目标函数,但该模型能够很好地进行对话。它能够从特定领域的数据集中以及从一个庞大、嘈杂且通用领域的电影字幕数据集中提取知识。在一个特定领域的IT服务台数据集上,该模型可以通过对话找到技术问题的解决方案。在一个嘈杂的开放领域电影脚本数据集上,该模型可以进行简单形式的常识推理。正如预期的那样,我们还发现缺乏一致性是我们模型的一种常见失效模式。
Conversational modeling is an important task in natural language understanding and machine intelligence. Although previous approaches exist, they are often restricted to specific domains (e.g., booking an airline ticket) and require hand-crafted rules. In this paper, we present a simple approach for this task which uses the recently proposed sequence to sequence framework. Our model converses by predicting the next sentence given the previous sentence or sentences in a conversation. The strength of our model is that it can be trained end-to-end and thus requires much fewer hand-crafted rules. We find that this straightforward model can generate simple conversations given a large conversational training dataset. Our preliminary results suggest that, despite optimizing the wrong objective function, the model is able to converse well. It is able extract knowledge from both a domain specific dataset, and from a large, noisy, and general domain dataset of movie subtitles. On a domain-specific IT helpdesk dataset, the model can find a solution to a technical problem via conversations. On a noisy open-domain movie transcript dataset, the model can perform simple forms of common sense reasoning. As expected, we also find that the lack of consistency is a common failure mode of our model.