Fine-Grained Action Retrieval Through Multiple Parts-of-Speech Embeddings

Fine-Grained Action Retrieval Through Multiple Parts-of-Speech Embeddings
复制标题

DOI:
10.1109/iccv.2019.00054
复制
发表时间:
2019-08
期刊:
2019 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
通讯作者:
Michael Wray;Diane Larlus;G. Csurka;D. Damen
Michael Wray;Diane Larlus;G. Csurka;D. Damen
中科院分区:
其他
文献类型:
--
作者:
Michael Wray;Diane Larlus;G. Csurka;D. Damen

文献摘要

被引文献

相似文献

我们解决了文本和视频之间的跨模态细粒度动作检索的问题。跨模态检索通常通过学习共享的嵌入空间来实现,该嵌入空间可以无差别地嵌入模态。在本文中,我们建议通过在附带的字幕中解开词性(PoS)来丰富嵌入。我们为每个PoS标签构建一个单独的多模态嵌入空间。然后,多个PoS嵌入的输出被用作集成多模态空间的输入,在那里我们执行动作检索。所有嵌入都通过PoS感知和PoS不可知损失的组合进行联合训练。我们的建议使学习专门的嵌入空间,提供相同的嵌入实体的多个视图。我们报告的第一个检索结果细粒度的行动,为大规模的EPIC数据集,在一个广义的零杆设置。结果表明,我们的方法的视频到文本和文本到视频动作检索的优势。我们还展示了在MSR-VTT数据集上为跨模态视频检索的通用任务解开PoS的好处。
We address the problem of cross-modal fine-grained action retrieval between text and video. Cross-modal retrieval is commonly achieved through learning a shared embedding space, that can indifferently embed modalities. In this paper, we propose to enrich the embedding by disentangling parts-of-speech (PoS) in the accompanying captions. We build a separate multi-modal embedding space for each PoS tag. The outputs of multiple PoS embeddings are then used as input to an integrated multi-modal space, where we perform action retrieval. All embeddings are trained jointly through a combination of PoS-aware and PoS-agnostic losses. Our proposal enables learning specialised embedding spaces that offer multiple views of the same embedded entities. We report the first retrieval results on fine-grained actions for the large-scale EPIC dataset, in a generalised zero-shot setting. Results show the advantage of our approach for both video-to-text and text-to-video action retrieval. We also demonstrate the benefit of disentangling the PoS for the generic task of cross-modal video retrieval on the MSR-VTT dataset.