Analyzing the Contribution of Commonsense Knowledge Sources for Why-Question Answering

Analyzing the Contribution of Commonsense Knowledge Sources for Why-Question Answering
复制标题

DOI:
--
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Yash Kumar Lal
Yash Kumar Lal
中科院分区:
其他
文献类型:
--
作者:
Yash Kumar Lal

文献摘要

相似文献

回答关于故事中事件发生的原因的问题需要常识性知识,而常识性知识是叙事之外的。大型模型可以访问003的哪些方面?哪些方面可以通过外部常识性资源获得?我们在回答007 TellMeWhy数据集中的Why问题的背景下研究了这些问题,使用COMET作为相关常识关系的来源。我们分析了在(a)增加模型大小,(b)将COMET的知识作为任务输入的一部分注入,以及(c)要求模型生成彗星关系类型作为其答案之外的解释014时,相对于基本T5 010模型的相对改进。结果表明,正如预期的那样,015更大的模型在基础上产生了实质性的016改进。有趣的是,我们发现特定于问题的COMET关系可以为基础模型和大型模型提供实质性的改进,当要求模型也生成021 COMET关系类型时,还可能获得额外的收益。因此,我们用来自COMET的噪声提示增强了一个大型的022模型,并发现这提高了TellMe-024 Why任务的性能。我们还建立了一个包含025种知识类型的简单本体,并分析了不同模型在这些类别上的相对覆盖范围。总之,这些发现表明,028方法有可能自动从相关来源选择和注入029常识。030
Answering questions about why events happen 001 in narratives requires commonsense knowledge 002 that is external to the narrative. What aspects of 003 this knowledge is accessible to large models? 004 What aspects can be made accessible via exter-005 nal commonsense resources? We study these in 006 the context of answering Why questions in the 007 TellMeWhy dataset using COMET as a source 008 of relevant commonsense relations. We ana-009 lyze the relative improvements over a base T5 010 model when (a) increasing the model size, (b) 011 injecting knowledge from COMET as part of 012 the task input, and (c) asking the model to gen-013 erate COMET relation type as an explanation 014 in addition to its answer. Results show that the 015 larger model, as expected, yields substantial 016 improvements over the base. Interestingly, we 017 find that the question specific COMET relations 018 can provide substantial improvements for both 019 base and large models, with additional possible 020 gains when asking the model to also generate 021 COMET relation type. So, we augment a large 022 model with noisy hints from COMET and find 023 that this improves performance on the TellMe-024 Why task. We also develop a simple ontology of 025 knowledge types and analyze the relative cover-026 age of the different models on these categories. 027 Together, these findings suggest potential for 028 methods that can automatically select and inject 029 commonsense from relevant sources. 030