Incorporating visual features into word embeddings: A bimodal autoencoder-based approach

Incorporating visual features into word embeddings: A bimodal autoencoder-based approach
复制标题

DOI:
--
复制
发表时间:
2017
期刊:
--
影响因子:
--
通讯作者:
Mika Hasegawa;Tetsunori Kobayashi;Yoshihiko Hayashi
Mika Hasegawa;Tetsunori Kobayashi;Yoshihiko Hayashi
中科院分区:
其他
文献类型:
--
作者:
Mika Hasegawa;Tetsunori Kobayashi;Yoshihiko Hayashi

文献摘要

被引文献

相似文献

多模态语义表示是自然语言处理和计算机视觉中一个不断发展的研究领域。将视觉特征等感知信息与语言特征相结合或整合是近年来研究的热点。本文提出了一种新的用于多模态表示学习的双模态自编码器模型:自动编码器通过结合相应的视觉特征来学习以增强语言特征向量。在运行期间,由于训练的神经网络,即使对于没有学习到直接视觉语言对应的单词,也可以实现视觉增强的多模态表示。通过标准语义关联任务获得的经验结果表明,我们的方法总体上是有希望的。我们进一步研究了增强词嵌入在区分反义词和同义词和模糊相关词方面的潜在功效。
Multimodal semantic representation is an evolving area of research in natural language processing as well as computer vision. Combining or integrating perceptual information, such as visual features, with linguistic features is recently being actively studied. This paper presents a novel bi-modal autoencoder model for multimodal representation learning: the autoencoder learns in order to enhance linguistic feature vectors by incorporating the corresponding visual features. During the runtime, owing to the trained neural network, visually enhanced multimodal representations can be achieved even for words for which direct visual-linguistic correspondences are not learned. The empirical results obtained with standard semantic relatedness tasks demonstrate that our approach is generally promising. We further investigate the potential efficacy of the enhanced word embeddings in discriminating antonyms and synonyms from vaguely related words.