Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?

Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?
复制标题

DOI:
10.18653/v1/2022.emnlp-main.759
复制
发表时间:
2022-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Sewon Min;Xinxi Lyu;Ari Holtzman;Mikel Artetxe;M. Lewis;Hannaneh Hajishirzi;Luke Zettlemoyer
Sewon Min;Xinxi Lyu;Ari Holtzman;Mikel Artetxe;M. Lewis;Hannaneh Hajishirzi;Luke Zettlemoyer
中科院分区:
其他
文献类型:
--
作者:
Sewon Min;Xinxi Lyu;Ari Holtzman;Mikel Artetxe;M. Lewis;Hannaneh Hajishirzi;Luke Zettlemoyer

文献摘要

相似文献

大型语言模型(LM)能够在上下文中学习-仅通过推理来执行新任务,通过对一些输入-标签对(演示)进行调节并对新输入进行预测。然而,人们对模型如何学习以及演示的哪些方面有助于最终任务性能的理解很少。在本文中,我们表明,地面真相演示实际上是不需要的随机替换标签的演示几乎不会损害性能的一系列分类和多选择任务,一致超过12个不同的模型,包括GPT-3。相反,我们发现演示的其他方面是最终任务性能的关键驱动因素,包括它们提供了以下几个示例:(1)标签空间,(2)输入文本的分布,以及(3)序列的整体格式。总之,我们的分析提供了一种新的方式来理解上下文学习是如何以及为什么工作的,同时也提出了一些新的问题,即仅仅通过推理就可以从大型语言模型中学到多少东西。
Large language models (LMs) are able to in-context learn—perform a new task via inference alone by conditioning on a few input-label pairs (demonstrations) and making predictions for new inputs. However, there has been little understanding of how the model learns and which aspects of the demonstrations contribute to end task performance. In this paper, we show that ground truth demonstrations are in fact not required—randomly replacing labels in the demonstrations barely hurts performance on a range of classification and multi-choce tasks, consistently over 12 different models including GPT-3. Instead, we find that other aspects of the demonstrations are the key drivers of endtask performance, including the fact that they provide a few examples of (1) the label space, (2) the distribution of the input text, and (3) the overall format of the sequence. Together, our analysis provides a new way of understanding how and why in-context learning works, while opening up new questions about how much can be learned from large language models through inference alone.