Tiered Reasoning for Intuitive Physics: Toward Verifiable Commonsense Language Understanding

Tiered Reasoning for Intuitive Physics: Toward Verifiable Commonsense Language Understanding
复制标题

DOI:
10.18653/v1/2021.findings-emnlp.422
复制
发表时间:
2021-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Shane Storks;Qiaozi Gao;Yichi Zhang;J. Chai
Shane Storks;Qiaozi Gao;Yichi Zhang;J. Chai
中科院分区:
其他
文献类型:
--
作者:
Shane Storks;Qiaozi Gao;Yichi Zhang;J. Chai

文献摘要

被引文献

相似文献

大规模的预训练语言模型(LM)在广泛的语言理解任务上取得了人类水平的表现。然而,仅基于最终任务表现的评估对机器在语言理解和推理方面的真正能力几乎没有帮助。在本文中,我们强调的重要性,评估的基础推理过程中,除了最终性能。为了实现这一目标,我们引入了直观物理分层推理(TRIP),这是一种具有密集注释的新型常识推理数据集,可以对机器的推理过程进行多层评估。我们的实证结果表明,虽然大型LM可以实现高端性能,但它们很难用有效的支持证据来支持它们的预测。TRIP数据集和我们的基线结果将激励对常识推理的可验证评估,并促进未来的研究,以开发更好的语言理解和推理模型。
Large-scale, pre-trained language models (LMs) have achieved human-level performance on a breadth of language understanding tasks. However, evaluations only based on end task performance shed little light on machines' true ability in language understanding and reasoning. In this paper, we highlight the importance of evaluating the underlying reasoning process in addition to end performance. Toward this goal, we introduce Tiered Reasoning for Intuitive Physics (TRIP), a novel commonsense reasoning dataset with dense annotations that enable multi-tiered evaluation of machines' reasoning process. Our empirical results show that while large LMs can achieve high end performance, they struggle to support their predictions with valid supporting evidence. The TRIP dataset and our baseline results will motivate verifiable evaluation of commonsense reasoning and facilitate future research toward developing better language understanding and reasoning models.