Using Commonsense Knowledge to Answer Why-Questions

Using Commonsense Knowledge to Answer Why-Questions
复制标题

DOI:
10.18653/v1/2022.emnlp-main.79
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Yash Kumar Lal
Yash Kumar Lal
中科院分区:
其他
文献类型:
--
作者:
Yash Kumar Lal

文献摘要

相似文献

在叙述中提出为什么事件发生的问题通常需要文本之外的常识知识。这些知识的哪些方面在大型语言模型中可用?哪些方面可以通过外部常识资源进行访问?我们在使用COMET作为相关常识关系来源的TellMeWhy数据集中回答问题的背景下研究这些问题。我们分析了模型大小(T5和GPT3)的影响,沿着注入知识的方法(COMET)到这些模型。结果表明,正如预期的那样,最大的模型比基本模型有很大的改进。注入外部知识有助于不同规模的模型,但随着模型规模的增大,改进的数量会减少。我们还发现,提供知识的格式是至关重要的,更小的模型受益于更大量的知识。最后,我们开发了一个知识类型的本体,并分析了这些类别的模型的相对覆盖率。
Answering questions in narratives about why events happened often requires commonsense knowledge external to the text. What aspects of this knowledge are available in large language models? What aspects can be made accessible via external commonsense resources? We study these questions in the context of answering questions in the TellMeWhy dataset using COMET as a source of relevant commonsense relations. We analyze the effects of model size (T5 and GPT3) along with methods of injecting knowledge (COMET) into these models. Results show that the largest models, as expected, yield substantial improvements over base models. Injecting external knowledge helps models of various sizes, but the amount of improvement decreases with larger model size. We also find that the format in which knowledge is provided is critical, and that smaller models benefit more from larger amounts of knowledge. Finally, we develop an ontology of knowledge types and analyze the relative coverage of the models across these categories.