Incorporating visual features into word embeddings: A bimodal autoencoder-based approach
Incorporating visual features into word embeddings: A bimodal autoencoder-based approach
复制标题
DOI:
--
复制
发表时间:
2017
期刊:
影响因子:
--
通讯作者:
Mika Hasegawa;Tetsunori Kobayashi;Yoshihiko Hayashi
中科院分区:
文献类型:
--
作者:
Mika Hasegawa;Tetsunori Kobayashi;Yoshihiko Hayashi
Multimodal semantic representation is an evolving area of research in natural language processing as well as computer vision. Combining or integrating perceptual information, such as visual features, with linguistic features is recently being actively studied. This paper presents a novel bi-modal autoencoder model for multimodal representation learning: the autoencoder learns in order to enhance linguistic feature vectors by incorporating the corresponding visual features. During the runtime, owing to the trained neural network, visually enhanced multimodal representations can be achieved even for words for which direct visual-linguistic correspondences are not learned. The empirical results obtained with standard semantic relatedness tasks demonstrate that our approach is generally promising. We further investigate the potential efficacy of the enhanced word embeddings in discriminating antonyms and synonyms from vaguely related words.