What Can Transformers Learn In-Context? A Case Study of Simple Function Classes

What Can Transformers Learn In-Context? A Case Study of Simple Function Classes
复制标题

DOI:
10.48550/arxiv.2208.01066
复制
发表时间:
2022-08
期刊:
ArXiv
影响因子:
--
通讯作者:
Shivam Garg;Dimitris Tsipras;Percy Liang;G. Valiant
Shivam Garg;Dimitris Tsipras;Percy Liang;G. Valiant
中科院分区:
其他
文献类型:
--
作者:
Shivam Garg;Dimitris Tsipras;Percy Liang;G. Valiant

文献摘要

被引文献

相似文献

中文学习是指模型在及时示例(输入输出对对应于某些任务的输入输出对)以及新的查询输入以及生成相应输出的及时示例(输入输出对对应)的能力。至关重要的是,在推理时间仅发生任何参数更新,因此内部文化学习仅在推理时间发生。尽管大型语言模型(例如GPT-3)具有某种能力来执行中文学习的能力,但尚不清楚任务成功的任务与培训数据中存在的内容之间的关系。为了取得进步朝着理解文本学习的进步,我们考虑了训练模型的明确定义的问题,以学习函数类(例如,线性函数):即,给定的数据从类中的某些功能衍生而成,可以我们训练一个模型以在此课程中学习“大多数”功能?我们从经验上表明,可以从头开始训练标准的变压器,以执行线性功能的秘密学习 - 也就是说,训练有素的模型能够从具有与最佳最小二乘估计器相当的性能的内在示例中学习看不见的线性函数。实际上,即使在两种形式的分布变化下,也可能进行中文学习:(i)模型的训练数据与推理时间提示之间,以及(ii)在推理过程中的示例和查询输入之间。我们还表明,我们可以训练变压器以学习更多复杂的功能类,即稀疏线性功能,两层神经网络和决策树 - 具有匹配或超过特定于任务特定的学习算法的性能。我们的代码和模型可在https://github.com/dtsip/in-context-learning上找到。
In-context learning refers to the ability of a model to condition on a prompt sequence consisting of in-context examples (input-output pairs corresponding to some task) along with a new query input, and generate the corresponding output. Crucially, in-context learning happens only at inference time without any parameter updates to the model. While large language models such as GPT-3 exhibit some ability to perform in-context learning, it is unclear what the relationship is between tasks on which this succeeds and what is present in the training data. To make progress towards understanding in-context learning, we consider the well-defined problem of training a model to in-context learn a function class (e.g., linear functions): that is, given data derived from some functions in the class, can we train a model to in-context learn"most"functions from this class? We show empirically that standard Transformers can be trained from scratch to perform in-context learning of linear functions -- that is, the trained model is able to learn unseen linear functions from in-context examples with performance comparable to the optimal least squares estimator. In fact, in-context learning is possible even under two forms of distribution shift: (i) between the training data of the model and inference-time prompts, and (ii) between the in-context examples and the query input during inference. We also show that we can train Transformers to in-context learn more complex function classes -- namely sparse linear functions, two-layer neural networks, and decision trees -- with performance that matches or exceeds task-specific learning algorithms. Our code and models are available at https://github.com/dtsip/in-context-learning .