Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning

Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning
复制标题

DOI:
10.48550/arxiv.2205.05638
复制
发表时间:
2022-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Haokun Liu;Derek Tam;Mohammed Muqeeth;Jay Mohta;Tenghao Huang;Mohit Bansal;Colin Raffel
Haokun Liu;Derek Tam;Mohammed Muqeeth;Jay Mohta;Tenghao Huang;Mohit Bansal;Colin Raffel
中科院分区:
其他
文献类型:
--
作者:
Haokun Liu;Derek Tam;Mohammed Muqeeth;Jay Mohta;Tenghao Huang;Mohit Bansal;Colin Raffel

文献摘要

相似文献

少镜头上下文学习(ICL)通过提供少量的训练样本作为输入的一部分,使预先训练的语言模型能够执行以前未见过的任务,而无需任何基于梯度的训练。ICL会产生大量的计算、内存和存储成本,因为它涉及在每次做出预测时处理所有训练样本。参数高效微调(PEFT)(例如适配器模块、提示调优、稀疏更新方法等)提供了一种替代范例,其中训练了一小部分参数以使模型能够执行新任务。在这篇文章中,我们严格地比较了几次ICL和PEFT,并证明了后者提供了更好的精度以及显著降低的计算成本。在此过程中,我们引入了一种新的PEFT方法,称为(IA)$^3$,它通过学习向量来扩展激活,在只引入相对少量的新参数的情况下获得更强的性能。我们还提出了一个基于T0模型的简单配方,称为T-少数,它可以应用于新任务,而不需要特定于任务的调优或修改。我们通过将其应用于RAFT基准测试,首次获得超人性能,并以6%的绝对值超过最先进的水平,从而验证了T-少数在完全看不见的任务上的有效性。我们实验中使用的所有代码都是公开可用的。
Few-shot in-context learning (ICL) enables pre-trained language models to perform a previously-unseen task without any gradient-based training by feeding a small number of training examples as part of the input. ICL incurs substantial computational, memory, and storage costs because it involves processing all of the training examples every time a prediction is made. Parameter-efficient fine-tuning (PEFT) (e.g. adapter modules, prompt tuning, sparse update methods, etc.) offers an alternative paradigm where a small set of parameters are trained to enable a model to perform the new task. In this paper, we rigorously compare few-shot ICL and PEFT and demonstrate that the latter offers better accuracy as well as dramatically lower computational costs. Along the way, we introduce a new PEFT method called (IA)$^3$ that scales activations by learned vectors, attaining stronger performance while only introducing a relatively tiny amount of new parameters. We also propose a simple recipe based on the T0 model called T-Few that can be applied to new tasks without task-specific tuning or modifications. We validate the effectiveness of T-Few on completely unseen tasks by applying it to the RAFT benchmark, attaining super-human performance for the first time and outperforming the state-of-the-art by 6% absolute. All of the code used in our experiments is publicly available.