How Much Knowledge Can You Pack into the Parameters of a Language Model?

How Much Knowledge Can You Pack into the Parameters of a Language Model?
复制标题

DOI:
10.18653/v1/2020.emnlp-main.437
复制
发表时间:
2020-02
期刊:
--
影响因子:
--
通讯作者:
Adam Roberts;Colin Raffel;Noam M. Shazeer
Adam Roberts;Colin Raffel;Noam M. Shazeer
中科院分区:
其他
文献类型:
--
作者:
Adam Roberts;Colin Raffel;Noam M. Shazeer

文献摘要

被引文献

相似文献

最近观察到,在非结构化文本上训练的神经语言模型可以使用自然语言查询隐式地存储和检索知识。在这篇简短的论文中,我们通过微调预训练模型来衡量这种方法的实际效用,以在不访问任何外部背景或知识的情况下回答问题。我们发现,这种方法的规模与模型大小惊人的好,并优于模型,明确查找知识的开放域的自然问题和网络问题的变体。为了促进可重复性和未来的工作,我们发布了我们的代码和训练模型。
It has recently been observed that neural language models trained on unstructured text can implicitly store and retrieve knowledge using natural language queries. In this short paper, we measure the practical utility of this approach by fine-tuning pre-trained models to answer questions without access to any external context or knowledge. We show that this approach scales surprisingly well with model size and outperforms models that explicitly look up knowledge on the open-domain variants of Natural Questions and WebQuestions. To facilitate reproducibility and future work, we release our code and trained models.