Multimodal Deep Embedding via Hierarchical Grounded Compositional Semantics

Multimodal Deep Embedding via Hierarchical Grounded Compositional Semantics
复制标题

DOI:
10.1109/tcsvt.2016.2606648
复制
发表时间:
2018
影响因子:
8.4
通讯作者:
Yueting Zhuang;Jun Song;Fei Wu;Xi Li;Zhongfei Zhang;Y. Rui
Yueting Zhuang;Jun Song;Fei Wu;Xi Li;Zhongfei Zhang;Y. Rui
中科院分区:
工程技术1区
文献类型:
--
作者:
Yueting Zhuang;Jun Song;Fei Wu;Xi Li;Zhongfei Zhang;Y. Rui

文献摘要

被引文献

相似文献

对于许多重要的问题,单独的句法单词或视觉对象的孤立语义表示是不够的,而是需要组合语义表示;例如,字面短语或一组空间上并发的对象。在本文中,我们的目标是利用现有的图像句子数据库来开发图像句子数据的组成性质,用于多模式深度嵌入。特别是,我们提出了一种称为类层次(自下而上两层)的多通道扎根组合语义(HiMoCS)学习方法。HiMoCS通过对文本实体(及其描述属性)和视觉对象、短语(如主谓宾语三元组)和空间并发对象之间的内在关联进行建模,系统地捕捉了分层深度学习环境下多通道数据的组成语义内涵。我们认为,HiMoCS更适合反映图像及其叙述性语篇句子的强耦合的多通道构成语义。我们在几个基准数据集上对hiMoCS进行了评估,结果表明,使用hiMoCS(文本实体和视觉对象、文本短语和空间并发对象)比只使用平坦的组合语义获得了更好的性能。
For a number of important problems, isolated semantic representations of individual syntactic words or visual objects do not suffice, but instead a compositional semantic representation is required; for example, a literal phrase or a set of spatially concurrent objects. In this paper, we aim to harness the existing image–sentence databases to exploit the compositional nature of image–sentence data for multimodal deep embedding. In particular, we propose an approach called hierarchical-alike (bottom–up two layers) multimodal grounded compositional semantics (hiMoCS) learning. The proposed hiMoCS systemically captures the compositional semantic connotation of multimodal data in the setting of hierarchical-alike deep learning by modeling the inherent correlations between two modalities of collaboratively grounded semantics, such as the textual entity (with its describing attribute) and visual object, the phrase (e.g., subject-verb–object triplet), and spatially concurrent objects. We argue that hiMoCS is more appropriate to reflect the multimodal compositional semantics of the image and its narrative textual sentence, which are strongly coupled. We evaluate hiMoCS on the several benchmark data sets and show that the utilization of the hiMoCS (textual entities and visual objects, textual phrase, and spatially concurrent objects) achieves a much better performance than only using the flat grounded compositional semantics.