Understanding and predicting user dissatisfaction in a neural generative chatbot

Understanding and predicting user dissatisfaction in a neural generative chatbot
复制标题

理解和预测神经生成聊天机器人中的用户不满

DOI:
10.18653/v1/2021.sigdial-1.1
复制
发表时间:
2021
期刊:
ArXiv
影响因子:
--
通讯作者:
A. See
A. See
中科院分区:
--
文献类型:
--
作者:
A. See

文献摘要

被引文献

相似文献

神经生成对话代理已经显示出越来越多的能力来举行短暂的闲聊对话,当众包工作者在受控环境中进行评估时。然而,它们在实际部署中的表现--在嘈杂环境中与具有内在动机的用户交谈--却没有得到很好的研究。在这篇文章中,我们对一个神经生成模型进行了详细的案例研究,该模型是作为Alexa奖社交机器人Chirpy Cardinal的一部分部署的。我们发现,不清晰的用户话语是产生忽视、幻觉、不清楚和重复等生成性错误的主要来源。然而,即使在明确的上下文中,该模型也经常犯推理错误。尽管用户对这些错误表达了相关的不满,但某些不满类型(如冒犯和隐私异议)取决于其他因素--如用户的个人态度,以及之前在对话中未解决的不满。最后,我们证明了不满意的用户话语可以作为半监督学习信号来改进对话系统。我们训练了一个模型来预测下一轮的不满,并通过人类评价表明,作为一个排序函数,它选择了更高质量的神经生成话语。
Neural generative dialogue agents have shown an increasing ability to hold short chitchat conversations, when evaluated by crowdworkers in controlled settings. However, their performance in real-life deployment – talking to intrinsically-motivated users in noisy environments – is less well-explored. In this paper, we perform a detailed case study of a neural generative model deployed as part of Chirpy Cardinal, an Alexa Prize socialbot. We find that unclear user utterances are a major source of generative errors such as ignoring, hallucination, unclearness and repetition. However, even in unambiguous contexts the model frequently makes reasoning errors. Though users express dissatisfaction in correlation with these errors, certain dissatisfaction types (such as offensiveness and privacy objections) depend on additional factors – such as the user’s personal attitudes, and prior unaddressed dissatisfaction in the conversation. Finally, we show that dissatisfied user utterances can be used as a semi-supervised learning signal to improve the dialogue system. We train a model to predict next-turn dissatisfaction, and show through human evaluation that as a ranking function, it selects higher-quality neural-generated utterances.