Learning New Tricks From Old Dogs: Multi-Source Transfer Learning From Pre-Trained Networks

Learning New Tricks From Old Dogs: Multi-Source Transfer Learning From Pre-Trained Networks
复制标题

DOI:
--
复制
发表时间:
2019
期刊:
--
影响因子:
--
通讯作者:
Joshua K. Lee;P. Sattigeri;G. Wornell
Joshua K. Lee;P. Sattigeri;G. Wornell
中科院分区:
其他
文献类型:
--
作者:
Joshua K. Lee;P. Sattigeri;G. Wornell

文献摘要

相似文献

用于移动的设备和传感器的深度学习算法的出现,导致了在广泛的机器学习任务上训练的系统的可用性和数量的急剧扩展,在迁移学习领域创造了大量的机遇和挑战。目前,大多数迁移学习方法都需要对学习到的系统进行某种控制,要么在源训练期间强制执行约束,要么在任务之间使用联合优化目标,要求所有数据都位于同一位置进行训练。然而,出于实际、隐私或其他原因,在各种应用中,我们可能无法控制单个源任务训练,也无法访问源训练样本。相反,我们只能访问在这些数据上预先训练的特征,作为“黑盒”的输出。对于这样的场景,我们考虑多源学习问题,即使用预先训练的神经网络的集合来训练分类器,用于一组未被任何源网络观察到的类,并且我们只有很少的训练样本。我们表明,通过使用这些分布式网络作为特征提取器,我们可以使用(非线性)最大相关分析工具以计算效率高的方式训练有效的分类器。特别是,我们开发了一种称为最大相关加权(MCW)的方法,用于从源网络的特征函数的适当加权中构建所需的目标分类器。我们说明了所得分类器对来自CIFAR-100,斯坦福大学Dogs和Tiny ImageNet数据集的数据集的有效性,此外,使用该方法来表征不同源任务在学习目标任务时的相对价值。
The advent of deep learning algorithms for mobile devices and sensors has led to a dramatic expansion in the availability and number of systems trained on a wide range of machine learning tasks, creating a host of opportunities and challenges in the realm of transfer learning. Currently, most transfer learning methods require some kind of control over the systems learned, either by enforcing constraints during the source training, or through the use of a joint optimization objective between tasks that requires all data be co-located for training. However, for practical, privacy, or other reasons, in a variety of applications we may have no control over the individual source task training, nor access to source training samples. Instead we only have access to features pre-trained on such data as the output of "black-boxes.'' For such scenarios, we consider the multi-source learning problem of training a classifier using an ensemble of pre-trained neural networks for a set of classes that have not been observed by any of the source networks, and for which we have very few training samples. We show that by using these distributed networks as feature extractors, we can train an effective classifier in a computationally-efficient manner using tools from (nonlinear) maximal correlation analysis. In particular, we develop a method we refer to as maximal correlation weighting (MCW) to build the required target classifier from an appropriate weighting of the feature functions from the source networks. We illustrate the effectiveness of the resulting classifier on datasets derived from the CIFAR-100, Stanford Dogs, and Tiny ImageNet datasets, and, in addition, use the methodology to characterize the relative value of different source tasks in learning a target task.