DELTA: DEep Learning Transfer using Feature Map with Attention for Convolutional Networks

DELTA: DEep Learning Transfer using Feature Map with Attention for Convolutional Networks
复制标题

DOI:
--
复制
发表时间:
2019-01
期刊:
ArXiv
影响因子:
--
通讯作者:
Xingjian Li;Haoyi Xiong;Hanchao Wang;Yuxuan Rao;Liping Liu;Jun Huan
Xingjian Li;Haoyi Xiong;Hanchao Wang;Yuxuan Rao;Liping Liu;Jun Huan
中科院分区:
其他
文献类型:
--
作者:
Xingjian Li;Haoyi Xiong;Hanchao Wang;Yuxuan Rao;Liping Liu;Jun Huan

文献摘要

被引文献

相似文献

通过微调具有超大数据集的预训练神经网络(如ImageNet)来进行迁移学习,可以显着加快训练速度,而准确性经常受到新目标任务有限数据集大小的影响。为了解决这个问题,人们研究了一些以起点为参考(SPAR)约束目标网络外层权重的正则化方法。在本文中,我们提出了一种新的正则化迁移学习框架DELTA,即使用注意力特征映射的深度学习迁移。DELTA算法不对神经网络的权值进行约束,而是保持目标网络的外层输出。具体来说,除了最小化经验损失外,DELTA还打算通过约束由以监督学习方式学习的注意力精确选择的特征映射子集来对齐两个网络的外层输出。实验结果表明,该方法在新任务中具有更高的准确性,优于这些基线。
Transfer learning through fine-tuning a pre-trained neural network with an extremely large dataset, such as ImageNet, can significantly accelerate training while the accuracy is frequently bottlenecked by the limited dataset size of the new target task. To solve the problem, some regularization methods, constraining the outer layer weights of the target network using the starting point as references (SPAR), have been studied. In this paper, we propose a novel regularized transfer learning framework DELTA, namely DEep Learning Transfer using Feature Map with Attention. Instead of constraining the weights of neural network, DELTA aims to preserve the outer layer outputs of the target network. Specifically, in addition to minimizing the empirical loss, DELTA intends to align the outer layer outputs of two networks, through constraining a subset of feature maps that are precisely selected by attention that has been learned in an supervised learning manner. We evaluate DELTA with the state-of-the-art algorithms, including L2 and L2-SP. The experiment results show that our proposed method outperforms these baselines with higher accuracy for new tasks.