An Efficient Policy Gradient Method for Conditional Dialogue Generation

An Efficient Policy Gradient Method for Conditional Dialogue Generation
复制标题

DOI:
10.1109/icdm.2019.00013
复制
发表时间:
2019-11
期刊:
2019 IEEE International Conference on Data Mining (ICDM)
影响因子:
--
通讯作者:
Lei Cai;Shuiwang Ji
Lei Cai;Shuiwang Ji
中科院分区:
其他
文献类型:
--
作者:
Lei Cai;Shuiwang Ji

文献摘要

相似文献

编码器-解码器模型已广泛用于对话生成任务。然而,他们往往会产生枯燥且笼统的回应话语。为了解决这个问题,我们将对话生成视为条件生成问题。对于给定的上下文历史,我们的模型可以生成具有所需对话行为的不同响应话语。我们的模型遵循 SeqGAN 框架,其中生成器将上下文历史记录和对话行为作为输入并生成相应的响应话语。判别器通过考虑整个话语和对话行为的质量来计算奖励。我们的模型是通过策略梯度方法进行训练的。为了克服蒙特卡罗搜索训练所带来的时间复杂度过高的瓶颈,我们提出了一种局部判别器网络来计算一次前向传播中的个体奖励,从而大大加速了训练过程。实验结果表明,我们提出的方法可以实现与蒙特卡洛搜索相当的性能,同时显着减少训练时间。
Encoder-decoder models have been commonly used in dialogue generation tasks. However, they tend to generate dull and generic response utterances. To tackle this problem, we consider dialogue generation as a conditional generation problem. For a given context history, our model can generate different response utterances with desirable dialog acts. Our model follows the SeqGAN framework, where the generator takes context history and dialog act as inputs and generates corresponding response utterances. The discriminator computes rewards by considering the quality of entire utterance and dialog act. Our model is trained by a policy gradient approach. To overcome the bottleneck of excessive time complexity incurred by the Monte Carlo search for training, we propose a local discriminator network to compute the individual reward in one forward propagation, thereby dramatically accelerating the training procedure. Experimental results demonstrate that our proposed method can achieve comparative performance with Monte Carlo search, while reducing the training time dramatically.