Making Scalable Meta Learning Practical

Making Scalable Meta Learning Practical
复制标题

DOI:
10.48550/arxiv.2310.05674
复制
发表时间:
2023-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Sang Keun Choe;Sanket Vaibhav Mehta;Hwijeen Ahn;W. Neiswanger;Pengtao Xie;Emma Strubell;Eric P. Xing
Sang Keun Choe;Sanket Vaibhav Mehta;Hwijeen Ahn;W. Neiswanger;Pengtao Xie;Emma Strubell;Eric P. Xing
中科院分区:
其他
文献类型:
--
作者:
Sang Keun Choe;Sanket Vaibhav Mehta;Hwijeen Ahn;W. Neiswanger;Pengtao Xie;Emma Strubell;Eric P. Xing

文献摘要

相似文献

尽管元学习在机器学习程序中可以灵活地学习不同的归纳偏差,但由于其巨大的计算/存储成本、训练的不稳定性以及缺乏有效的分布式训练支持,元学习(即学习学习)长期以来一直被认为存在可扩展性差的问题。在这项工作中,我们专注于通过引入SAMA来使可伸缩的元学习变得实用,SAMA结合了隐式微分算法和系统的优点。具体地说,SAMA被设计成在元学习程序的基础水平上灵活地支持广泛的自适应优化器,同时通过避免显式计算二阶梯度信息和利用针对一阶梯度实现的高效分布式训练技术来减少计算负担。在多个大规模元学习基准上进行评估,SAMA显示与其他基准元学习算法相比,在单/多GPU设置上,吞吐量提高了1.7/4.8倍,内存消耗减少了2.0/3.8倍。此外,基于SAMA的数据优化导致了BERT和Roberta大语言模型在文本分类准确率方面的持续提高,并在图像分类任务的小规模和大规模数据剪枝方面取得了最先进的结果,展示了跨语言和视觉领域的可伸缩元学习的实用适用性。
Despite its flexibility to learn diverse inductive biases in machine learning programs, meta learning (i.e., learning to learn) has long been recognized to suffer from poor scalability due to its tremendous compute/memory costs, training instability, and a lack of efficient distributed training support. In this work, we focus on making scalable meta learning practical by introducing SAMA, which combines advances in both implicit differentiation algorithms and systems. Specifically, SAMA is designed to flexibly support a broad range of adaptive optimizers in the base level of meta learning programs, while reducing computational burden by avoiding explicit computation of second-order gradient information, and exploiting efficient distributed training techniques implemented for first-order gradients. Evaluated on multiple large-scale meta learning benchmarks, SAMA showcases up to 1.7/4.8x increase in throughput and 2.0/3.8x decrease in memory consumption respectively on single-/multi-GPU setups compared to other baseline meta learning algorithms. Furthermore, we show that SAMA-based data optimization leads to consistent improvements in text classification accuracy with BERT and RoBERTa large language models, and achieves state-of-the-art results in both small- and large-scale data pruning on image classification tasks, demonstrating the practical applicability of scalable meta learning across language and vision domains.