Describing Unseen Videos via Multi-Modal Cooperative Dialog Agents
Describing Unseen Videos via Multi-Modal Cooperative Dialog Agents
复制标题
DOI:
10.1007/978-3-030-58592-1_10
复制
发表时间:
2020-08
期刊:
影响因子:
--
通讯作者:
Ye Zhu;Yu Wu;Yi Yang;Yan Yan-Yan
中科院分区:
文献类型:
--
作者:
Ye Zhu;Yu Wu;Yi Yang;Yan Yan-Yan
With the arising concerns for the AI systems provided with direct access to abundant sensitive information, researchers seek to develop more reliable AI with implicit information sources. To this end, in this paper, we introduce a new task called video description via two multi-modal cooperative dialog agents, whose ultimate goal is for one conversational agent to describe an unseen video based on the dialog and two static frames. Specifically, one of the intelligent agents -Q-BOT- is given two static frames from the beginning and the end of the video, as well as a finite number of opportunities to ask relevant natural language questions before describing the unseen video.A-BOT, the other agent who has already seen the entire video, assistsQ-BOTto accomplish the goal by providing answers to those questions. We propose a QA-Cooperative Network with a dynamic dialog history update learning mechanism to transfer knowledge fromA-BOTtoQ-BOT, thus helpingQ-BOTto better describe the video. Extensive experiments demonstrate thatQ-BOTcan effectively learn to describe an unseen video by the proposed model and the cooperative learning method, achieving the promising performance whereQ-BOTis given the full ground truth history dialog. Codes and models are available at https://github.com/L-YeZhu/Video-Description-via-Dialog-Agents-ECCV2020 .