Unnatural Instructions: Tuning Language Models with (Almost) No Human Labor

Unnatural Instructions: Tuning Language Models with (Almost) No Human Labor
复制标题

DOI:
10.48550/arxiv.2212.09689
复制
发表时间:
2022-12
期刊:
ArXiv
影响因子:
--
通讯作者:
Or Honovich;Thomas Scialom;Omer Levy;Timo Schick
Or Honovich;Thomas Scialom;Omer Levy;Timo Schick
中科院分区:
其他
文献类型:
--
作者:
Or Honovich;Thomas Scialom;Omer Levy;Timo Schick

文献摘要

被引文献

相似文献

指令调优使预训练的语言模型能够从推理时的自然语言描述执行新任务。这些方法依赖于以众包数据集或用户交互形式进行的大量人工监督。在这项工作中,我们介绍了非自然指令:一个创造性和多样化指令的大型数据集,几乎没有人工劳动。我们收集了64,000个例子,通过提示一个语言模型,其中包含三个指令的种子例子,并引出第四个。然后通过提示模型重新表述每个指令来扩展该集合,创建总计约240,000个指令、输入和输出示例。实验表明,尽管包含了相当数量的噪声,但在非自然指令上的训练可以与在开源人工管理数据集上的训练相媲美,在各种基准测试中超过了T0++和tk - directive等模型的性能。这些结果证明了模型生成数据作为一种具有成本效益的替代方法的潜力,可以用于数据集扩展和多样化。
Instruction tuning enables pretrained language models to perform new tasks from inference-time natural language descriptions. These approaches rely on vast amounts of human supervision in the form of crowdsourced datasets or user interactions. In this work, we introduce Unnatural Instructions: a large dataset of creative and diverse instructions, collected with virtually no human labor. We collect 64,000 examples by prompting a language model with three seed examples of instructions and eliciting a fourth. This set is then expanded by prompting the model to rephrase each instruction, creating a total of approximately 240,000 examples of instructions, inputs, and outputs. Experiments show that despite containing a fair amount of noise, training on Unnatural Instructions rivals the effectiveness of training on open-source manually-curated datasets, surpassing the performance of models such as T0++ and Tk-Instruct across various benchmarks. These results demonstrate the potential of model-generated data as a cost-effective alternative to crowdsourcing for dataset expansion and diversification.