Cross-Modality Feature Learning via Convolutional Autoencoder

Cross-Modality Feature Learning via Convolutional Autoencoder
复制标题

通过卷积自动编码器进行跨模态特征学习

DOI:
10.1145/3231740
复制
发表时间:
2019
期刊:
ACM Transactions on Multimedia Computing, Communications, and Applications
影响因子:
--
通讯作者:
Hong Richang
Hong Richang
中科院分区:
其他
文献类型:
--
作者:
Liu Xueliang;Wang Meng;Zha Zheng-Jun;Hong Richang

文献摘要

被引文献

相似文献

学习跨多模态的鲁棒性和代表性特征一直是机器学习和多媒体领域的基本问题。在本文中,我们提出了一种新的多模态卷积自动编码器(MUCAE)方法来从视觉和文本模式中学习代表性特征。对于每种模态,我们将卷积操作集成到自动编码器框架中,以从原始图像和文本内容中学习联合表示。我们通过利用卷积自编码器隐藏表征之间的相关性来共同优化不同模态的卷积自编码器,特别是通过最小化每个模态的重构误差和不同模态隐藏特征之间的相关性分歧来实现。与依赖手工特征的传统解决方案相比,本文提出的MUCAE方法直接从图像像素和文本字符中编码特征,产生更具代表性和鲁棒性的特征。我们评估了MUCAE在跨媒体检索以及在真实世界的大型多媒体数据库上的单模分类任务。实验结果表明,MUCAE比最先进的方法具有更好的性能。
Learning robust and representative features across multiple modalities has been a fundamental problem in machine learning and multimedia fields. In this article, we propose a novel MUltimodal Convolutional AutoEncoder (MUCAE) approach to learn representative features from visual and textual modalities. For each modality, we integrate the convolutional operation into an autoencoder framework to learn a joint representation from the original image and text content. We optimize the convolutional autoencoders of different modalities jointly by exploiting the correlation between the hidden representations from the convolutional autoencoders, in particular by minimizing both the reconstructing error of each modality and the correlation divergence between the hidden feature of different modalities. Compared to the conventional solutions relying on hand-crafted features, the proposed MUCAE approach encodes features from image pixels and text characters directly and produces more representative and robust features. We evaluate MUCAE on cross-media retrieval as well as unimodal classification tasks over real-world large-scale multimedia databases. Experimental results have shown that MUCAE performs better than the state-of-the-arts methods.