Distill or Annotate? Cost-Efficient Fine-Tuning of Compact Models

Distill or Annotate? Cost-Efficient Fine-Tuning of Compact Models
复制标题

DOI:
10.48550/arxiv.2305.01645
复制
发表时间:
2023-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Junmo Kang;Wei Xu;Alan Ritter
Junmo Kang;Wei Xu;Alan Ritter
中科院分区:
其他
文献类型:
--
作者:
Junmo Kang;Wei Xu;Alan Ritter

文献摘要

相似文献

微调大型模型非常有效,但是推理成本可能很高并且会产生碳排放。知识蒸馏已被证明是降低推理成本的实用解决方案,但蒸馏过程本身需要大量的计算资源。 NLP 从业者可能会选择分配可用预算来雇用注释者并手动标记额外的微调数据,而不是购买或租用 GPU 进行微调,然后提取大型模型。在本文中,我们研究如何最有效地使用固定预算来构建紧凑的模型。通过对六种不同任务的广泛实验,我们表明,与注释更多数据以直接训练紧凑模型(T5-Small)相比,从 T5-XXL(11B)提炼到 T5-Small(60M)几乎总是一种经济高效的策略。我们进一步研究分配给计算的最佳预算如何随场景而变化。我们将提供我们的代码、数据集、注释成本估算和基线模型作为基准,以支持紧凑模型的成本效益训练的进一步工作。
Fine-tuning large models is highly effective, however, inference can be expensive and produces carbon emissions. Knowledge distillation has been shown to be a practical solution to reduce inference costs, but the distillation process itself requires significant computational resources. Rather than buying or renting GPUs to fine-tune, then distill a large model, an NLP practitioner might instead choose to allocate the available budget to hire annotators and manually label additional fine-tuning data. In this paper, we investigate how to most efficiently use a fixed budget to build a compact model. Through extensive experiments on six diverse tasks, we show that distilling from T5-XXL (11B) to T5-Small (60M) is almost always a cost-efficient strategy compared to annotating more data to directly train a compact model (T5-Small). We further investigate how the optimal budget allocated towards computation varies across scenarios. We will make our code, datasets, annotation cost estimates, and baseline models available as a benchmark to support further work on cost-efficient training of compact models.