Decoding Visual Neural Representations by Multimodal Learning of Brain-Visual-Linguistic Features

Decoding Visual Neural Representations by Multimodal Learning of Brain-Visual-Linguistic Features
复制标题

DOI:
10.1109/tpami.2023.3263181
复制
发表时间:
2023-09-01
影响因子:
23.6
通讯作者:
He, Huiguang
He, Huiguang
中科院分区:
计算机科学1区
文献类型:
--
作者:
Du, Changde;Fu, Kaicheng;He, Huiguang

文献摘要

被引文献

相似文献

解码人类视觉神经表征是一项具有挑战性的任务,对于揭示视觉处理机制和开发类脑智能机器具有重要的科学意义。大多数现有的方法很难推广到没有相应的神经数据进行训练的新类别。主要有两个原因:1)神经数据背后的多模态语义知识开发不足;2)配对(刺激反应)训练数据数量少。为了克服这些限制,本文提出了一种称为BraVL的通用神经解码方法,该方法使用脑-视觉-语言特征的多模态学习。我们专注于通过多模态深度生成模型来建模大脑、视觉和语言特征之间的关系。具体地说,我们利用专家的产品混合公式来推断一个潜在的代码,使所有三种模式的连贯联合生成。为了在有限的大脑活动数据情况下学习更一致的联合表示并提高数据效率,我们利用了模态内和模态间的互信息最大化正则化项。特别是,我们的BraVL模型可以在各种半监督场景下进行训练,以结合从额外类别中获得的视觉和文本特征。最后,我们构建了三个三模态匹配数据集,并通过大量的实验得出了一些有趣的结论和认知见解:1)从人类大脑活动中解码新的视觉类别实际上是可能的,并且具有良好的准确性;2)视觉特征和语言特征相结合的解码模型比单独使用其中任何一种的解码模型表现更好;3)视觉知觉可能伴随着语言影响来表征视觉刺激的语义。
Decoding human visual neural representations is a challenging task with great scientific significance in revealing vision-processing mechanisms and developing brain-like intelligent machines. Most existing methods are difficult to generalize to novel categories that have no corresponding neural data for training. The two main reasons are 1) the under-exploitation of the multimodal semantic knowledge underlying the neural data and 2) the small number of paired (stimuli-responses) training data. To overcome these limitations, this paper presents a generic neural decoding method called BraVL that uses multimodal learning of brain-visual-linguistic features. We focus on modeling the relationships between brain, visual and linguistic features via multimodal deep generative models. Specifically, we leverage the mixture-of-product-of-experts formulation to infer a latent code that enables a coherent joint generation of all three modalities. To learn a more consistent joint representation and improve the data efficiency in the case of limited brain activity data, we exploit both intra- and inter-modality mutual information maximization regularization terms. In particular, our BraVL model can be trained under various semi-supervised scenarios to incorporate the visual and textual features obtained from the extra categories. Finally, we construct three trimodal matching datasets, and the extensive experiments lead to some interesting conclusions and cognitive insights: 1) decoding novel visual categories from human brain activity is practically possible with good accuracy; 2) decoding models using the combination of visual and linguistic features perform much better than those using either of them alone; 3) visual perception may be accompanied by linguistic influences to represent the semantics of visual stimuli.