Learning Multimodal Representations for Unseen Activities

Learning Multimodal Representations for Unseen Activities
复制标题

DOI:
10.1109/wacv45572.2020.9093612
复制
发表时间:
2018-06
期刊:
2020 IEEE Winter Conference on Applications of Computer Vision (WACV)
影响因子:
--
通讯作者:
A. Piergiovanni;M. Ryoo
A. Piergiovanni;M. Ryoo
中科院分区:
其他
文献类型:
--
作者:
A. Piergiovanni;M. Ryoo

文献摘要

相似文献

我们提出了一种学习联合多模态表示空间的方法,该空间能够识别视频中看不见的活动。我们首先使用成对的文本和视频数据来比较在嵌入空间上放置各种约束的效果。我们还提出了一种使用对抗公式来改进联合嵌入空间的方法,使其能够从未配对的文本和视频数据中受益。通过使用未配对的文本数据,我们展示了学习更好地捕捉看不见的活动的表示的能力。除了在公开数据集上进行测试外,我们还引入了一个新的大规模文本/视频数据集。我们通过实验证实,使用配对和未配对的数据来学习共享嵌入空间有利于三个困难的任务:(i)零触发活动分类,(ii)无监督活动发现,以及(iii)看不见的活动字幕,优于最先进的技术。
We present a method to learn a joint multimodal representation space that enables recognition of unseen activities in videos. We first compare the effect of placing various constraints on the embedding space using paired text and video data. We also propose a method to improve the joint embedding space using an adversarial formulation, allowing it to benefit from unpaired text and video data. By using unpaired text data, we show the ability to learn a representation that better captures unseen activities. In addition to testing on publicly available datasets, we introduce a new, large-scale text/video dataset. We experimentally confirm that using paired and unpaired data to learn a shared embedding space benefits three difficult tasks (i) zero-shot activity classification, (ii) unsupervised activity discovery, and (iii) unseen activity captioning, outperforming the state-of-the-arts.