History for Visual Dialog: Do we really need it?

History for Visual Dialog: Do we really need it?
复制标题

DOI:
10.18653/v1/2020.acl-main.728
复制
发表时间:
2020-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Shubham Agarwal;Trung Bui;Joon-Young Lee;Ioannis Konstas;Verena Rieser
Shubham Agarwal;Trung Bui;Joon-Young Lee;Ioannis Konstas;Verena Rieser
中科院分区:
其他
文献类型:
--
作者:
Shubham Agarwal;Trung Bui;Joon-Young Lee;Ioannis Konstas;Verena Rieser

文献摘要

被引文献

相似文献

视觉对话涉及“理解”对话历史(前面已经讨论过的内容)和当前问题(被问到的内容),以及图像中的基础信息,以准确地生成正确的响应。在本文中,我们证明了显式编码dialoh历史的共同注意模型优于不编码的模型,实现了最先进的性能(瓦尔集上72%的NDCG)。然而,我们也暴露了众包数据集收集过程的缺点,通过显示对话历史确实只需要少量的数据,并且当前的评估指标鼓励通用的回复。为此,我们提出了一个具有挑战性的子集(VisdialConv)的VisdialVal集和基准NDCG的63%。
Visual Dialogue involves “understanding” the dialogue history (what has been discussed previously) and the current question (what is asked), in addition to grounding information in the image, to accurately generate the correct response. In this paper, we show that co-attention models which explicitly encode dialoh history outperform models that don’t, achieving state-of-the-art performance (72 % NDCG on val set). However, we also expose shortcomings of the crowdsourcing dataset collection procedure, by showing that dialogue history is indeed only required for a small amount of the data, and that the current evaluation metric encourages generic replies. To that end, we propose a challenging subset (VisdialConv) of the VisdialVal set and the benchmark NDCG of 63%.