When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories

When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories
复制标题

DOI:
10.18653/v1/2023.acl-long.546
复制
发表时间:
2022-12
期刊:
--
影响因子:
--
通讯作者:
Alex Troy Mallen;Akari Asai;Victor Zhong;Rajarshi Das;Hannaneh Hajishirzi;Daniel Khashabi
Alex Troy Mallen;Akari Asai;Victor Zhong;Rajarshi Das;Hannaneh Hajishirzi;Daniel Khashabi
中科院分区:
其他
文献类型:
--
作者:
Alex Troy Mallen;Akari Asai;Victor Zhong;Rajarshi Das;Hannaneh Hajishirzi;Daniel Khashabi

文献摘要

被引文献

相似文献

尽管大型语言模型(LM)在不同的任务上表现出色,但它们仍然难以处理需要丰富世界知识的任务,这意味着在其参数中编码丰富的世界知识是困难的。本文旨在了解LM在记忆事实知识方面的优势和局限性,通过对两个开放域以实体为中心的QA数据集进行大规模知识探测实验:PopQA,我们的新数据集,包含14k个关于长尾实体的问题,以及一个广泛使用的开放域QA数据集。我们发现,LM斗争较少流行的事实知识,检索增强在这些情况下有很大帮助。另一方面,缩放主要提高了大众知识的记忆,而未能明显提高尾部事实知识的记忆。基于这些研究结果,我们设计了一种新的方法检索增强,提高性能,降低推理成本,只有检索非参数记忆时,必要的。
Despite their impressive performance on diverse tasks, large language models (LMs) still struggle with tasks requiring rich world knowledge, implying the difficulty of encoding a wealth of world knowledge in their parameters. This paper aims to understand LMs’ strengths and limitations in memorizing factual knowledge, by conducting large-scale knowledge probing experiments on two open-domain entity-centric QA datasets: PopQA, our new dataset with 14k questions about long-tail entities, and EntityQuestions, a widely used open-domain QA dataset. We find that LMs struggle with less popular factual knowledge, and that retrieval augmentation helps significantly in these cases. Scaling, on the other hand, mainly improves memorization of popular knowledge, and fails to appreciably improve memorization of factual knowledge in the tail. Based on those findings, we devise a new method for retrieval-augmentation that improves performance and reduces inference costs by only retrieving non-parametric memories when necessary.