X-CNN: Cross-modal convolutional neural networks for sparse datasets

X-CNN: Cross-modal convolutional neural networks for sparse datasets
复制标题

DOI:
10.1109/ssci.2016.7849978
复制
发表时间:
2016-10
期刊:
2016 IEEE Symposium Series on Computational Intelligence (SSCI)
影响因子:
--
通讯作者:
Petar Velickovic;Duo Wang;N. Lane;P. Lio’
Petar Velickovic;Duo Wang;N. Lane;P. Lio’
中科院分区:
其他
文献类型:
--
作者:
Petar Velickovic;Duo Wang;N. Lane;P. Lio’

文献摘要

相似文献

在本文中,我们提出了交叉模态卷积神经网络(X-CNN),这是一种新型的生物学启发型CNN架构,将梯度下降专用CNN视为大规模网络拓扑中的单个处理单元,同时允许网络的类似隐藏层之间的不受约束的信息流和/或权重共享(其中信息通常仅在各个网络的输出层之间流动)。组成网络被单独设计为在它们自己的输入数据子集上学习输出函数,之后在每次池化操作之后引入它们之间的交叉连接,以定期允许它们之间的信息交换。这种将知识注入模型(通过领域知识或无监督方法对输入数据进行先验划分)预计将在稀疏数据环境中产生最大的回报,而稀疏数据环境通常不太适合训练CNN。出于评估目的,我们将标准的四层CNN以及复杂的FitNet 4架构与CIFAR-10和CIFAR-100数据集上的跨模态变体进行了比较,其中删除了不同百分比的训练数据,并发现在较低的数据可用性水平下,X-CNN显着优于其基线。(通常提供2-6%的收益,取决于数据集大小和是否使用数据增强),同时仍然保持所有完整数据集测试的优势。
In this paper we propose cross-modal convolutional neural networks (X-CNNs), a novel biologically inspired type of CNN architectures, treating gradient descent-specialised CNNs as individual units of processing in a larger-scale network topology, while allowing for unconstrained information flow and/or weight sharing between analogous hidden layers of the network—thus generalising the already well-established concept of neural network ensembles (where information typically may flow only between the output layers of the individual networks). The constituent networks are individually designed to learn the output function on their own subset of the input data, after which cross-connections between them are introduced after each pooling operation to periodically allow for information exchange between them. This injection of knowledge into a model (by prior partition of the input data through domain knowledge or unsupervised methods) is expected to yield greatest returns in sparse data environments, which are typically less suitable for training CNNs. For evaluation purposes, we have compared a standard four-layer CNN as well as a sophisticated FitNet4 architecture against their cross-modal variants on the CIFAR-10 and CIFAR-100 datasets with differing percentages of the training data being removed, and find that at lower levels of data availability, the X-CNNs significantly outperform their baselines (typically providing a 2–6% benefit, depending on the dataset size and whether data augmentation is used), while still maintaining an edge on all of the full dataset tests.