Cross-Modal Prediction of Superclasses Using Cortex-Inspired Neural Architecture

Cross-Modal Prediction of Superclasses Using Cortex-Inspired Neural Architecture
复制标题

DOI:
10.1109/biosmart58455.2023.10162118
复制
发表时间:
2023-06
期刊:
2023 5th International Conference on Bio-engineering for Smart Technologies (BioSMART)
影响因子:
--
通讯作者:
Olcay Kursun;Hoa T. Nguyen;O. Favorov
Olcay Kursun;Hoa T. Nguyen;O. Favorov
中科院分区:
其他
文献类型:
--
作者:
Olcay Kursun;Hoa T. Nguyen;O. Favorov

文献摘要

相似文献

刺激特征调谐的概念是神经科学的基础。皮质神经元通过从经验中学习并使用来自这些特征发生的空间和/或时间背景的暂定特征的潜在有用性的代理符号来获得其特征调谐特性。根据这一思想,局部但最终在行为上有用的特征应该是那些与其他此类特征可预测相关的特征,这些特征要么在时间上先于它们,要么与它们并排发生。受这一想法的启发,本文将深度神经网络与典型相关分析(CCA)相结合进行特征提取,并使用无监督跨模态预测任务展示了特征的强大功能。CCA是一种多视图特征提取方法,可以在多个数据集(通常称为视图或模态)中找到相关特征。CCA发现每个视图的线性变换,使得提取的主成分或特征具有最大的互相关性。CCA是一种线性方法,特征是通过每个视图变量的加权和来计算的。一旦学习了权重,CCA就可以应用于新示例,并通过从源(查询)视图中的给定变量推断示例的目标视图特征来用于跨模态预测。为了测试所提出的方法,它被应用于非结构化的CIFAR-100数据集的60,000图像分类成100类,进一步分为20个超类,并用于演示挖掘图像标签的相关性。对三个预先训练的CNN的输出进行CCA:AlexNet,ResNet和VGG。利用CCA提取的互相关特征,在查询视图和目标视图共有的规范子空间中搜索最近邻,在目标视图中检索最匹配的示例,无需任何监督训练即可成功预测测试视图的超类成员。
The concept of stimulus feature tuning is fundamental to neuroscience. Cortical neurons acquire their feature-tuning properties by learning from experience and using proxy signs of tentative features’ potential usefulness that come from the spatial and/or temporal context in which these features occur. According to this idea, local but ultimately behaviorally useful features should be the ones that are predictably related to other such features either preceding them in time or taking place side-by-side with them. Inspired by this idea, in this paper, deep neural networks are combined with Canonical Correlation Analysis (CCA) for feature extraction and the power of the features is demonstrated using unsupervised cross-modal prediction tasks. CCA is a multi-view feature extraction method that finds correlated features across multiple datasets (usually referred to as views or modalities). CCA finds linear transformations of each view such that the extracted principal components, or features, have a maximal mutual correlation. CCA is a linear method, and the features are computed by a weighted sum of each view’s variables. Once the weights are learned, CCA can be applied to new examples and used for cross-modal prediction by inferring the target-view features of an example from its given variables in a source (query) view. To test the proposed method, it was applied to the unstructured CIFAR-100 dataset of 60,000 images categorized into 100 classes, which are further grouped into 20 superclasses and used to demonstrate the mining of image-tag correlations. CCA was performed on the outputs of three pre-trained CNNs: AlexNet, ResNet, and VGG. Taking advantage of the mutually correlated features extracted with CCA, a search for nearest neighbors was performed in the canonical subspace common to both the query and the target views to retrieve the most matching examples in the target view, which successfully predicted the superclass membership of the tested views without any supervised training.