ExplaGraphs: An Explanation Graph Generation Task for Structured Commonsense Reasoning

ExplaGraphs: An Explanation Graph Generation Task for Structured Commonsense Reasoning
复制标题

DOI:
10.18653/v1/2021.emnlp-main.609
复制
发表时间:
2021-04
期刊:
ArXiv
影响因子:
--
通讯作者:
Swarnadeep Saha;Prateek Yadav;Lisa Bauer;Mohit Bansal
Swarnadeep Saha;Prateek Yadav;Lisa Bauer;Mohit Bansal
中科院分区:
其他
文献类型:
--
作者:
Swarnadeep Saha;Prateek Yadav;Lisa Bauer;Mohit Bansal

文献摘要

相似文献

最近的常识推理任务在本质上通常是有区别的,其中模型回答特定上下文的多项选择题。区分性任务是有限的,因为它们无法充分评估模型的推理能力,并用基本常识知识解释预测。它们还允许这样的模型使用推理捷径,而不是“出于正确的原因而正确”。在这项工作中,我们提出了ExplaGraphs,一个新的生成和结构化的常识推理任务(和相关的数据集)的解释图生成的立场预测。具体来说,给定一个信念和一个论点,模型必须预测论点是支持还是反对信念,并生成一个常识增强图,作为预测立场的非平凡,完整和明确的解释。我们通过一个新的验证和细化图形收集框架,提高图形质量(高达90%),通过多轮的验证和细化收集解释图。我们的图中有79%包含具有不同结构和推理深度的外部常识节点。接下来,我们提出了一个多层次的评估框架,包括自动度量和人工评估,检查生成的图的结构和语义的正确性及其与地面实况图的匹配程度。最后,我们提出了几个结构化的,常识增强的,文本生成模型作为强有力的起点,这个解释图生成任务,并观察到,有一个很大的差距与人类的表现,从而鼓励未来的工作,这一新的挑战性任务。
Recent commonsense-reasoning tasks are typically discriminative in nature, where a model answers a multiple-choice question for a certain context. Discriminative tasks are limiting because they fail to adequately evaluate the model’s ability to reason and explain predictions with underlying commonsense knowledge. They also allow such models to use reasoning shortcuts and not be “right for the right reasons”. In this work, we present ExplaGraphs, a new generative and structured commonsense-reasoning task (and an associated dataset) of explanation graph generation for stance prediction. Specifically, given a belief and an argument, a model has to predict if the argument supports or counters the belief and also generate a commonsense-augmented graph that serves as non-trivial, complete, and unambiguous explanation for the predicted stance. We collect explanation graphs through a novel Create-Verify-And-Refine graph collection framework that improves the graph quality (up to 90%) via multiple rounds of verification and refinement. A significant 79% of our graphs contain external commonsense nodes with diverse structures and reasoning depths. Next, we propose a multi-level evaluation framework, consisting of automatic metrics and human evaluation, that check for the structural and semantic correctness of the generated graphs and their degree of match with ground-truth graphs. Finally, we present several structured, commonsense-augmented, and text generation models as strong starting points for this explanation graph generation task, and observe that there is a large gap with human performance, thereby encouraging future work for this new challenging task.