SemEval-2018 Task 12: The Argument Reasoning Comprehension Task

SemEval-2018 Task 12: The Argument Reasoning Comprehension Task
复制标题

DOI:
10.18653/v1/s18-1121
复制
发表时间:
2018-06
期刊:
--
影响因子:
--
通讯作者:
Ivan Habernal;Henning Wachsmuth;Iryna Gurevych;Benno Stein
Ivan Habernal;Henning Wachsmuth;Iryna Gurevych;Benno Stein
中科院分区:
其他
文献类型:
--
作者:
Ivan Habernal;Henning Wachsmuth;Iryna Gurevych;Benno Stein

文献摘要

被引文献

相似文献

一个自然语言论证由一个主张以及作为主张前提的理由组成。解释推理的理由通常是含蓄的,因为它从上下文和常识中是清楚的。这使得理解论点对人类来说很容易,但对机器来说很难。本文总结了第一个共享任务的论点推理理解。给定一个前提和一个声明沿着一些主题信息,目标是自动识别两个候选人之间的正确保证,这两个候选人是合理的,词汇上接近,但实际上意味着相反的声明。我们描述了我们为任务构建的1970个实例的数据集,并概述了参与的21种计算方法,其中大多数使用神经网络。结果揭示了任务的复杂性,许多方法几乎没有提高约0.5的随机精度。尽管如此,最好的观察精度(0.712)强调了识别权证的原则可行性。我们的分析表明,包含外部知识是推理理解的关键。
A natural language argument is composed of a claim as well as reasons given as premises for the claim. The warrant explaining the reasoning is usually left implicit, as it is clear from the context and common sense. This makes a comprehension of arguments easy for humans but hard for machines. This paper summarizes the first shared task on argument reasoning comprehension. Given a premise and a claim along with some topic information, the goal was to automatically identify the correct warrant among two candidates that are plausible and lexically close, but in fact imply opposite claims. We describe the dataset with 1970 instances that we built for the task, and we outline the 21 computational approaches that participated, most of which used neural networks. The results reveal the complexity of the task, with many approaches hardly improving over the random accuracy of about 0.5. Still, the best observed accuracy (0.712) underlines the principle feasibility of identifying warrants. Our analysis indicates that an inclusion of external knowledge is key to reasoning comprehension.