Describing Unseen Videos via Multi-Modal Cooperative Dialog Agents

Describing Unseen Videos via Multi-Modal Cooperative Dialog Agents
复制标题

DOI:
10.1007/978-3-030-58592-1_10
复制
发表时间:
2020-08
期刊:
--
影响因子:
--
通讯作者:
Ye Zhu;Yu Wu;Yi Yang;Yan Yan-Yan
Ye Zhu;Yu Wu;Yi Yang;Yan Yan-Yan
中科院分区:
其他
文献类型:
--
作者:
Ye Zhu;Yu Wu;Yi Yang;Yan Yan-Yan

文献摘要

相似文献

随着人们对直接访问大量敏感信息的人工智能系统的担忧日益增加,研究人员寻求开发具有隐式信息源的更可靠的人工智能。为此,在本文中,我们通过两个多模态协作对话代理引入了一种称为视频描述的新任务,其最终目标是让一个对话代理基于对话和两个静态帧来描述看不见的视频。具体来说,其中一个智能代理 Q-BOT 会从视频的开头和结尾获得两个静态帧,并在描述未看过的视频之前有有限数量的机会提出相关的自然语言问题。另一个已经看过整个视频的代理 A-BOT 通过提供这些问题的答案来协助 Q-BOT 完成目标。我们提出了一种具有动态对话历史更新学习机制的 QA 合作网络,将知识从 A-BOT 转移到 Q-BOT,从而帮助 Q-BOT 更好地描述视频。大量的实验表明,Q-BOT 可以通过所提出的模型和协作学习方法有效地学习描述未见过的视频,在 Q-BOT 给出完整的地面真实历史对话的情况下实现了有希望的性能。代码和模型可在 https://github.com/L-YeZhu/Video-Description-via-Dialog-Agents-ECCV2020 获取。
With the arising concerns for the AI systems provided with direct access to abundant sensitive information, researchers seek to develop more reliable AI with implicit information sources. To this end, in this paper, we introduce a new task called video description via two multi-modal cooperative dialog agents, whose ultimate goal is for one conversational agent to describe an unseen video based on the dialog and two static frames. Specifically, one of the intelligent agents -Q-BOT- is given two static frames from the beginning and the end of the video, as well as a finite number of opportunities to ask relevant natural language questions before describing the unseen video.A-BOT, the other agent who has already seen the entire video, assistsQ-BOTto accomplish the goal by providing answers to those questions. We propose a QA-Cooperative Network with a dynamic dialog history update learning mechanism to transfer knowledge fromA-BOTtoQ-BOT, thus helpingQ-BOTto better describe the video. Extensive experiments demonstrate thatQ-BOTcan effectively learn to describe an unseen video by the proposed model and the cooperative learning method, achieving the promising performance whereQ-BOTis given the full ground truth history dialog. Codes and models are available at https://github.com/L-YeZhu/Video-Description-via-Dialog-Agents-ECCV2020 .