Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks

Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks
复制标题

DOI:
--
复制
发表时间:
2017
期刊:
--
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

被引文献

相似文献

对于N-way,K-shot分类,使用NK个示例的批量大小来计算每个梯度。对于Omniglot,5路卷积和非卷积MAML模型均使用步长α = 0.4的1个梯度步长和32个任务的Meta批量进行训练。使用相同步长α = 0.4的3个梯度步长评价网络。使用步长α = 0.1的5个梯度步长对20向卷积MAML模型进行训练和评价。在训练期间,Meta批处理大小被设置为16个任务。对于MiniImagenet,两个模型均使用大小为α = 0.01的5个梯度步长进行训练,并在测试时使用10个梯度步长进行评估。在Ravi & Larochelle(2017)之后,每个类使用15个示例来评估更新后的元梯度。我们分别使用4个和2个任务的Meta批量进行1次和5次训练。所有模型都在单个NVIDIA Pascal Titan X GPU上进行了60000次迭代训练。
For N-way, K-shot classification, each gradient is computed using a batch size of NK examples. For Omniglot, the 5-way convolutional and non-convolutional MAML models were each trained with 1 gradient step with step size α = 0.4 and a meta batch-size of 32 tasks. The network was evaluated using 3 gradient steps with the same step size α = 0.4. The 20-way convolutional MAML model was trained and evaluated with 5 gradient steps with step size α = 0.1. During training, the meta batch-size was set to 16 tasks. For MiniImagenet, both models were trained using 5 gradient steps of size α = 0.01, and evaluated using 10 gradient steps at test time. Following Ravi & Larochelle (2017), 15 examples per class were used for evaluating the post-update meta-gradient. We used a meta batch-size of 4 and 2 tasks for 1-shot and 5-shot training respectively. All models were trained for 60000 iterations on a single NVIDIA Pascal Titan X GPU.