Improving Few-Shot Generalization by Exploring and Exploiting Auxiliary Data

Improving Few-Shot Generalization by Exploring and Exploiting Auxiliary Data
复制标题

DOI:
10.48550/arxiv.2302.00674
复制
发表时间:
2023-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Alon Albalak;Colin Raffel;William Yang Wang
Alon Albalak;Colin Raffel;William Yang Wang
中科院分区:
其他
文献类型:
--
作者:
Alon Albalak;Colin Raffel;William Yang Wang

文献摘要

被引文献

相似文献

少样本学习在许多现实世界的应用中是有价值的,但是学习一个可推广的模型而不过度拟合到少数标记的数据点是具有挑战性的。在这项工作中,我们专注于使用辅助数据的少次学习(FLAD),这是一种训练范式,它假设在少次学习期间可以访问辅助数据,希望能够提高泛化能力。以前的工作已经提出了用于混合辅助数据和目标数据的自动化方法,但这些方法通常随辅助数据集的数量线性(或更糟)扩展,限制了它们的实用性。在这项工作中,我们将FLAD与多臂强盗设置的核心探索-利用困境相关联,并推导出其计算复杂性与辅助数据集数量无关的算法,使我们能够扩展到比先前方法多100倍的辅助数据集。我们提出了两种算法-EXP 3-FLAD和UCB 1-FLAD -并将它们与先前的FLAD方法进行比较,无论是探索还是开发,发现探索和开发的结合是至关重要的。通过大量的实验,我们发现我们的方法比所有现有的FLAD方法的性能高出4%,并导致前30亿个参数语言模型的性能优于1750亿个参数GPT-3。总的来说,我们的工作表明,发现更好,更有效的FLAD混合策略可能会提供一条可行的道路,大大提高少数学习的泛化能力。
Few-shot learning is valuable in many real-world applications, but learning a generalizable model without overfitting to the few labeled datapoints is challenging. In this work, we focus on Few-shot Learning with Auxiliary Data (FLAD), a training paradigm that assumes access to auxiliary data during few-shot learning in hopes of improving generalization. Previous works have proposed automated methods for mixing auxiliary and target data, but these methods typically scale linearly (or worse) with the number of auxiliary datasets, limiting their practicality. In this work we relate FLAD to the explore-exploit dilemma that is central to the multi-armed bandit setting and derive algorithms whose computational complexity is independent of the number of auxiliary datasets, allowing us to scale to 100x more auxiliary datasets than prior methods. We propose two algorithms -- EXP3-FLAD and UCB1-FLAD -- and compare them with prior FLAD methods that either explore or exploit, finding that the combination of exploration and exploitation is crucial. Through extensive experimentation we find that our methods outperform all pre-existing FLAD methods by 4% and lead to the first 3 billion parameter language models that outperform the 175 billion parameter GPT-3. Overall, our work suggests that the discovery of better, more efficient mixing strategies for FLAD may provide a viable path towards substantially improving generalization in few-shot learning.