Explainable Semantic Space by Grounding Language to Vision with Cross-Modal Contrastive Learning

Explainable Semantic Space by Grounding Language to Vision with Cross-Modal Contrastive Learning
复制标题

DOI:
--
复制
发表时间:
2021-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Yizhen Zhang;Minkyu Choi;Kuan Han;Zhongming Liu
Yizhen Zhang;Minkyu Choi;Kuan Han;Zhongming Liu
中科院分区:
其他
文献类型:
--
作者:
Yizhen Zhang;Minkyu Choi;Kuan Han;Zhongming Liu

文献摘要

被引文献

相似文献

在自然语言处理中,大多数模型试图仅从文本中学习语义表示。学习的表示编码了分布语义,但无法连接到关于物理世界的任何知识。相比之下,人类通过感知和行动中的概念基础来学习语言,而大脑则为认知编码基础语义。受这一概念和视觉语言学习的最新研究成果的启发,我们设计了一个基于视觉语言学习的双流模型。该模型包括基于VGG的视频流和基于BERT的语言流。这两条流汇合成一个共同的表现空间。通过跨模式对比学习,该模型首先学习将视觉和语言表示与MS Coco数据集对齐。该模型还学习通过跨模式注意模块使用语言查询来检索视觉对象,并通过与视觉基因组数据集的双线性算子来推断检索到的对象之间的视觉关系。经过训练,该模型的语言流是一个独立的语言模型,能够在视觉上扎根的语义空间中嵌入概念。这个语义空间显示了可以用人类直觉和神经生物学知识解释的主要维度。该语义空间中的单词嵌入是人类定义的语义特征规范的预测,并被分成感知上不同的簇。此外,基于视觉的语言模型还允许基于视觉知识的构图语言理解和基于图像、文本或其组合的查询的多模式图像搜索。
In natural language processing, most models try to learn semantic representations merely from texts. The learned representations encode the distributional semantics but fail to connect to any knowledge about the physical world. In contrast, humans learn language by grounding concepts in perception and action and the brain encodes grounded semantics for cognition. Inspired by this notion and recent work in vision-language learning, we design a two-stream model for grounding language learning in vision. The model includes a VGG-based visual stream and a Bert-based language stream. The two streams merge into a joint representational space. Through cross-modal contrastive learning, the model first learns to align visual and language representations with the MS COCO dataset. The model further learns to retrieve visual objects with language queries through a cross-modal attention module and to infer the visual relations between the retrieved objects through a bilinear operator with the Visual Genome dataset. After training, the language stream of this model is a stand-alone language model capable of embedding concepts in a visually grounded semantic space. This semantic space manifests principal dimensions explainable with human intuition and neurobiological knowledge. Word embeddings in this semantic space are predictive of human-defined norms of semantic features and are segregated into perceptually distinctive clusters. Furthermore, the visually grounded language model also enables compositional language understanding based on visual knowledge and multimodal image search with queries based on images, texts, or their combinations.