Precise Task Formalization Matters in Winograd Schema Evaluations

Precise Task Formalization Matters in Winograd Schema Evaluations
复制标题

精确的任务形式化在 Winograd 模式评估中很重要

DOI:
10.18653/v1/2020.emnlp-main.664
复制
发表时间:
2020
期刊:
ArXiv
影响因子:
--
通讯作者:
Samuel R. Bowman
Samuel R. Bowman
中科院分区:
--
文献类型:
--
作者:
Haokun Liu;William Huang;Dhara Mungra;Samuel R. Bowman

文献摘要

被引文献

相似文献

Winograd Schema Challenge(WSC)是一个受人尊敬的英语常识推理基准测试,最近在SuperGLUE排行榜上的表现从偶然准确率飙升至89%,推理能力相应大幅提高的证据相对较少。我们假设,这种改进大部分来自于最近任务形式化的变化-输入规范,损失函数和预训练参数的重用-由数据集的用户组合,而不是预训练模型推理能力的改进。我们对两个Winograd Schema数据集进行了消融,这些数据集在此激增之前和之后使用的形式化之间进行了插值,并发现(i)将任务框定为多项选择将性能提高了2-6个点,以及(ii)几种额外的技术,包括重新使用预先训练的语言建模头,可以减轻模型对超参数的极端敏感性。我们敦促未来的基准创建者施加额外的结构,以最大限度地减少正式化决策对报告结果的影响。
Performance on the Winograd Schema Challenge (WSC), a respected English commonsense reasoning benchmark, recently rocketed from chance accuracy to 89% on the SuperGLUE leaderboard, with relatively little corroborating evidence of a correspondingly large improvement in reasoning ability. We hypothesize that much of this improvement comes from recent changes in task formalization---the combination of input specification, loss function, and reuse of pretrained parameters---by users of the dataset, rather than improvements in the pretrained model's reasoning ability. We perform an ablation on two Winograd Schema datasets that interpolates between the formalizations used before and after this surge, and find (i) framing the task as multiple choice improves performance by 2-6 points and (ii) several additional techniques, including the reuse of a pretrained language modeling head, can mitigate the model's extreme sensitivity to hyperparameters. We urge future benchmark creators to impose additional structure to minimize the impact of formalization decisions on reported results.