Learning How to Ask: Querying LMs with Mixtures of Soft Prompts

Learning How to Ask: Querying LMs with Mixtures of Soft Prompts
复制标题

DOI:
10.18653/v1/2021.naacl-main.410
复制
发表时间:
2021-04
期刊:
ArXiv
影响因子:
--
通讯作者:
Guanghui Qin;J. Eisner
Guanghui Qin;J. Eisner
中科院分区:
其他
文献类型:
--
作者:
Guanghui Qin;J. Eisner

文献摘要

被引文献

相似文献

自然语言提示最近已被用于诱导预训练语言模型执行其他人工智能任务,使用填空范式(佩特罗尼等人,2019年)或少量样本外推范式(布朗等人,2020年)。例如,语言模型从其训练语料库中保留了事实性知识,这些知识可以通过要求它们在句子提示中“填空”来提取。然而,这个提示来自哪里呢?我们探索通过梯度下降学习提示的想法——要么对从先前工作中获取的提示进行微调,要么从随机初始化开始。我们的提示由“软词”组成,即不一定是来自语言模型的词类型嵌入的连续向量。此外,对于每个任务,我们优化提示的混合,了解哪些提示最有效以及如何将它们组合。在多个英语语言模型和任务中,我们的方法大大优于先前的方法,表明语言模型中隐含的事实性知识此前被低估了。而且,引出这种知识的成本很低:随机初始化几乎和有信息的初始化一样好。
Natural-language prompts have recently been used to coax pretrained language models into performing other AI tasks, using a fill-in-the-blank paradigm (Petroni et al., 2019) or a few-shot extrapolation paradigm (Brown et al., 2020). For example, language models retain factual knowledge from their training corpora that can be extracted by asking them to “fill in the blank” in a sentential prompt. However, where does this prompt come from? We explore the idea of learning prompts by gradient descent—either fine-tuning prompts taken from previous work, or starting from random initialization. Our prompts consist of “soft words,” i.e., continuous vectors that are not necessarily word type embeddings from the language model. Furthermore, for each task, we optimize a mixture of prompts, learning which prompts are most effective and how to ensemble them. Across multiple English LMs and tasks, our approach hugely outperforms previous methods, showing that the implicit factual knowledge in language models was previously underestimated. Moreover, this knowledge is cheap to elicit: random initialization is nearly as good as informed initialization.