Saying the Unseen: Video Descriptions via Dialog Agents

Saying the Unseen: Video Descriptions via Dialog Agents
复制标题

DOI:
10.1109/tpami.2021.3093360
复制
发表时间:
2021-06
影响因子:
23.6
通讯作者:
Ye Zhu;Yu Wu;Yi Yang;Yan Yan-Yan
Ye Zhu;Yu Wu;Yi Yang;Yan Yan-Yan
中科院分区:
计算机科学1区
文献类型:
--
作者:
Ye Zhu;Yu Wu;Yi Yang;Yan Yan-Yan

文献摘要

相似文献

当前的视觉和语言任务通常采用完整的视觉数据(如原始图像或视频)作为输入,然而,实际场景中可能经常存在由于各种原因(如固定摄像机限制视野或出于安全考虑故意阻挡视觉)而导致部分视觉信息无法访问的情况。作为向更实际的应用场景迈进的一步,我们引入了一个新的任务,旨在使用两个智能体之间的自然语言对话来描述视频,作为给定不完整视觉数据的补充信息源。与大多数现有的视觉语言任务不同,人工智能系统可以完全访问图像或视频片段,这些图像或视频片段可能会揭示敏感信息,如可识别的人脸或声音,我们有意限制人工智能系统的视觉输入,并寻求更安全和透明的信息媒介,即自然语言对话,以补充缺失的视觉信息。具体来说,其中一个智能代理- Q-BOT -从视频的开始和结束获得两个语义分段帧,以及在描述未见过的视频之前提出相关自然语言问题的有限数量的机会。A-BOT是另一个可以访问整个视频的代理,通过回答Q-BOT提出的问题来协助Q-BOT完成目标。我们引入了两种不同的实验设置,其中一种是生成式(即智能体自由生成问题和答案),另一种是判别式(即智能体从候选人中选择问题和答案)内部对话生成过程。利用提出的统一qa -协作网络,实验证明了两个对话代理之间的知识转移过程,以及使用自然语言对话作为不完全隐式视觉的补充的有效性。
Current vision and language tasks usually take complete visual data (e.g., raw images or videos) as input, however, practical scenarios may often consist the situations where part of the visual information becomes inaccessible due to various reasons e.g., restricted view with fixed camera or intentional vision block for security concerns. As a step towards the more practical application scenarios, we introduce a novel task that aims to describe a video using the natural language dialog between two agents as a supplementary information source given incomplete visual data. Different from most existing vision-language tasks where AI systems have full access to images or video clips, which may reveal sensitive information such as recognizable human faces or voices, we intentionally limit the visual input for AI systems and seek a more secure and transparent information medium, i.e., the natural language dialog, to supplement the missing visual information. Specifically, one of the intelligent agents - Q-BOT - is given two semantic segmented frames from the beginning and the end of the video, as well as a finite number of opportunities to ask relevant natural language questions before describing the unseen video. A-BOT, the other agent who has access to the entire video, assists Q-BOT to accomplish the goal by answering the asked questions. We introduce two different experimental settings with either a generative (i.e., agents generate questions and answers freely) or a discriminative (i.e., agents select the questions and answers from candidates) internal dialog generation process. With the proposed unified QA-Cooperative networks, we experimentally demonstrate the knowledge transfer process between the two dialog agents and the effectiveness of using the natural language dialog as a supplement for incomplete implicit visions.