AvgOut: A Simple Output-Probability Measure to Eliminate Dull Responses

AvgOut: A Simple Output-Probability Measure to Eliminate Dull Responses
复制标题

DOI:
10.1609/aaai.v34i05.6378
复制
发表时间:
2020-01
期刊:
--
影响因子:
--
通讯作者:
Tong Niu;Mohit Bansal
Tong Niu;Mohit Bansal
中科院分区:
其他
文献类型:
--
作者:
Tong Niu;Mohit Bansal

文献摘要

相似文献

许多序列对序列对话模型倾向于生成安全的、不提供信息的响应。已经做出了各种有益的努力,试图消除这些问题。然而,这些方法要么在推理过程中改进解码算法,要么依赖手工制作的特征,要么使用复杂的模型。在我们的工作中,我们建立了对话模型,在没有任何特征工程的情况下,动态地感知哪些话语或标记是枯燥的。具体地说,我们从一个简单而有效的自动度量AvgOut开始,该度量计算在训练期间解码器端所有时间步长的平均输出概率分布。该度量直接估计更有可能生成哪些令牌,从而使其成为模型多样性的真实评估(即,对于不同的模型,令牌概率应该更均匀地分布,而不是在几个枯燥的令牌处达到峰值)。然后,我们利用这一新的衡量标准提出了三个在不失去相关性的情况下促进多样性的模型。第一个模型MinAvgOut通过每个批次的输出分布直接最大化多样性分数;第二个模型Label Fine-Tuning(LFT)在源序列前面加上一个根据多样性分数连续缩放的标签来控制多样性水平;第三个模型RL采用强化学习,并将多样性分数作为奖励信号。此外,我们通过结合MinAvgOut和RL的损失项,对混合模型进行了实验。所有四个模型在多样性和相关性上都大大超过了它们的基本LSTM-RNN模型,并且与竞争基线相当或更好(也通过人类评估进行了验证)。此外,我们的方法与基本模型是正交的,使它们可以作为未来其他新兴的更好的对话模型的补充。
Many sequence-to-sequence dialogue models tend to generate safe, uninformative responses. There have been various useful efforts on trying to eliminate them. However, these approaches either improve decoding algorithms during inference, rely on hand-crafted features, or employ complex models. In our work, we build dialogue models that are dynamically aware of what utterances or tokens are dull without any feature-engineering. Specifically, we start with a simple yet effective automatic metric, AvgOut, which calculates the average output probability distribution of all time steps on the decoder side during training. This metric directly estimates which tokens are more likely to be generated, thus making it a faithful evaluation of the model diversity (i.e., for diverse models, the token probabilities should be more evenly distributed rather than peaked at a few dull tokens). We then leverage this novel metric to propose three models that promote diversity without losing relevance. The first model, MinAvgOut, directly maximizes the diversity score through the output distributions of each batch; the second model, Label Fine-Tuning (LFT), prepends to the source sequence a label continuously scaled by the diversity score to control the diversity level; the third model, RL, adopts Reinforcement Learning and treats the diversity score as a reward signal. Moreover, we experiment with a hybrid model by combining the loss terms of MinAvgOut and RL. All four models outperform their base LSTM-RNN model on both diversity and relevance by a large margin, and are comparable to or better than competitive baselines (also verified via human evaluation). Moreover, our approaches are orthogonal to the base model, making them applicable as an add-on to other emerging better dialogue models in the future.